Datasets and models
Datasets, benchmarks, and model releases
Public datasets, benchmarks, and model checkpoints released by the laboratory.
| Resource | Project | Focus | Venue / year | Links |
|---|---|---|---|---|
| IndicTalk | — | Persona-based multilingual conversational corpus (1.3M+ dialogues; 9 Indic languages) | arXiv 2026 | arXivHugging Face |
| COMI-LINGUA | Code-mixed NLP | Expert-annotated Hindi–English code-mixing (LID, POS, NER, normalisation, translation) | Findings of EMNLP 2025 | PaperarXivGitHubHugging Face |
| SangrahaTox | — | Multimodal cultural toxicity benchmark | Hugging Face release | Hugging Face |
| Commentator | Code-mixed NLP | Code-mixed multilingual annotation framework | EMNLP 2024 demo | PaperGitHub |
| PythonSaga | — | Benchmark for code-generating LLMs | Findings of EMNLP 2024 | PaperarXivGitHubHugging Face |
| LEGOBench | — | Scientific leaderboard generation | Findings of EMNLP 2024 | PaperGitHub |
| TempUN | — | Numerical-temporal reasoning in LLMs | Findings of EMNLP 2024 | Paper |
| Eka-Eval | Project EKA | Evaluation suite for Indian-language LLMs | arXiv 2025 | arXivGitHub |
| Model Hubs | — | Reddit sentiment set for auditing model popularity vs. performance | ICWSM 2025 | PaperarXivGitHubHugging Face |
| Triveni | — | Indic multimodal fine-tuning corpus (English, Hindi, Hinglish captions) | Hugging Face release | Hugging Face |
| triveni-raw | — | Larger Triveni pretraining split (Vaani + Flickr30k; Hindi, English, Hinglish) | Hugging Face release | Hugging Face |
| PI-Indic-Align | — | Persona–instruction alignment benchmark across 12 Indian languages | arXiv 2026 | arXivHugging Face |
| Gurukul | — | Curriculum-aligned educational question–answer pairs from Indian school textbooks | Hugging Face release | Hugging Face |
| MUTANT | — | Multi-sentential code-mixed Hinglish | Findings of EACL 2023 | PaperHugging Face |
| MMT | — | Multilingual, multi-topic Indian social media | C3NLP 2023 | PaperGitHubHugging Face |
| HinGE | — | Generation and evaluation of code-mixed Hinglish | Eval4NLP 2021 | PaperarXivHugging Face |
| PHINC | — | Parallel Hinglish social-media corpus for MT | W-NUT @ EMNLP 2020 | PaperarXivHugging Face |
| PoliWAM | — | Political discussions on WhatsApp | W-NUT 2021 | PaperarXivHugging Face |
| LineEX | Scholarly document AI | Data extraction from scientific line charts | WACV 2023 | PaperGitHub |
| TabLeX | Scholarly document AI | Structure and content extraction from scientific tables | ICDAR 2021 | PaperarXiv |
| COMPARE | — | Comparison discussions in peer reviews | JCDL 2021 | PaperarXivGitHub |
Model releases
Open models
Weights on Hugging Face. Ganga models are listed on the Project Unity page.
| Model | Project | Description | Links |
|---|---|---|---|
| COMI-LINGUA-TN | Code-mixed NLP | Text normalisation on COMI-LINGUA | Hugging Face |
| COMI-LINGUA-MLI | Code-mixed NLP | Matrix-language identification on COMI-LINGUA | Hugging Face |
| COMI-LINGUA-POS | Code-mixed NLP | POS tagging on COMI-LINGUA | Hugging Face |
| COMI-LINGUA-MT | Code-mixed NLP | Code-mixed translation on COMI-LINGUA | Hugging Face |
| COMI-LINGUA-LID | Code-mixed NLP | Language identification on COMI-LINGUA (token classification) | Hugging Face |
| COMI-LINGUA-NER | Code-mixed NLP | Named-entity recognition on COMI-LINGUA | Hugging Face |
| Ansh-128k | — | Multilingual Indic tokenizer (128k vocabulary) | Hugging Face |
| Ansh-256k | — | Multilingual Indic tokenizer (256k vocabulary) | Hugging Face |
| Ansh-160k | — | Multilingual Indic tokenizer (160k vocabulary) | Hugging Face |
| Ganga-2-1B | Project Unity | Instruct-tuned follow-on of Ganga-1B | Hugging Face |
| Ganga-en-hi-1B | Project Unity | Ganga-1B fine-tuned for English–Hindi translation | Hugging Face |
| Ganga-1B | Project Unity | Hindi language model pretrained from scratch | Hugging Face |