Datasets and models

Datasets, benchmarks, and model releases

Public datasets, benchmarks, and model checkpoints released by the laboratory.

ResourceProjectFocusVenue / yearLinks
IndicTalk Persona-based multilingual conversational corpus (1.3M+ dialogues; 9 Indic languages) arXiv 2026 arXivHugging Face
COMI-LINGUA Code-mixed NLP Expert-annotated Hindi–English code-mixing (LID, POS, NER, normalisation, translation) Findings of EMNLP 2025 PaperarXivGitHubHugging Face
SangrahaTox Multimodal cultural toxicity benchmark Hugging Face release Hugging Face
Commentator Code-mixed NLP Code-mixed multilingual annotation framework EMNLP 2024 demo PaperGitHub
PythonSaga Benchmark for code-generating LLMs Findings of EMNLP 2024 PaperarXivGitHubHugging Face
LEGOBench Scientific leaderboard generation Findings of EMNLP 2024 PaperGitHub
TempUN Numerical-temporal reasoning in LLMs Findings of EMNLP 2024 Paper
Eka-Eval Project EKA Evaluation suite for Indian-language LLMs arXiv 2025 arXivGitHub
Model Hubs Reddit sentiment set for auditing model popularity vs. performance ICWSM 2025 PaperarXivGitHubHugging Face
Triveni Indic multimodal fine-tuning corpus (English, Hindi, Hinglish captions) Hugging Face release Hugging Face
triveni-raw Larger Triveni pretraining split (Vaani + Flickr30k; Hindi, English, Hinglish) Hugging Face release Hugging Face
PI-Indic-Align Persona–instruction alignment benchmark across 12 Indian languages arXiv 2026 arXivHugging Face
Gurukul Curriculum-aligned educational question–answer pairs from Indian school textbooks Hugging Face release Hugging Face
MUTANT Multi-sentential code-mixed Hinglish Findings of EACL 2023 PaperHugging Face
MMT Multilingual, multi-topic Indian social media C3NLP 2023 PaperGitHubHugging Face
HinGE Generation and evaluation of code-mixed Hinglish Eval4NLP 2021 PaperarXivHugging Face
PHINC Parallel Hinglish social-media corpus for MT W-NUT @ EMNLP 2020 PaperarXivHugging Face
PoliWAM Political discussions on WhatsApp W-NUT 2021 PaperarXivHugging Face
LineEX Scholarly document AI Data extraction from scientific line charts WACV 2023 PaperGitHub
TabLeX Scholarly document AI Structure and content extraction from scientific tables ICDAR 2021 PaperarXiv
COMPARE Comparison discussions in peer reviews JCDL 2021 PaperarXivGitHub

Model releases

Open models

Weights on Hugging Face. Ganga models are listed on the Project Unity page.

ModelProjectDescriptionLinks
COMI-LINGUA-TN Code-mixed NLP Text normalisation on COMI-LINGUA Hugging Face
COMI-LINGUA-MLI Code-mixed NLP Matrix-language identification on COMI-LINGUA Hugging Face
COMI-LINGUA-POS Code-mixed NLP POS tagging on COMI-LINGUA Hugging Face
COMI-LINGUA-MT Code-mixed NLP Code-mixed translation on COMI-LINGUA Hugging Face
COMI-LINGUA-LID Code-mixed NLP Language identification on COMI-LINGUA (token classification) Hugging Face
COMI-LINGUA-NER Code-mixed NLP Named-entity recognition on COMI-LINGUA Hugging Face
Ansh-128k Multilingual Indic tokenizer (128k vocabulary) Hugging Face
Ansh-256k Multilingual Indic tokenizer (256k vocabulary) Hugging Face
Ansh-160k Multilingual Indic tokenizer (160k vocabulary) Hugging Face
Ganga-2-1B Project Unity Instruct-tuned follow-on of Ganga-1B Hugging Face
Ganga-en-hi-1B Project Unity Ganga-1B fine-tuned for English–Hindi translation Hugging Face
Ganga-1B Project Unity Hindi language model pretrained from scratch Hugging Face