Dataset

COMI-LINGUA

Expert-annotated Hindi–English code-mixing (181k+ instances) for language identification, POS, NER, normalisation, and translation.

Paper and data

COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing

Sheth, Beniwal, Singh. Findings of EMNLP 2025.

PaperarXivGitHubHugging Face

Related

Task models

LID, POS, NER, MLI, MT, and TN models trained on this dataset.

Models page →

Theme →Datasets and models →Publications →