COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
Sheth, Beniwal, Singh. Findings of EMNLP 2025.
PaperarXivGitHubHugging FaceDataset
Expert-annotated Hindi–English code-mixing (181k+ instances) for language identification, POS, NER, normalisation, and translation.
Sheth, Beniwal, Singh. Findings of EMNLP 2025.
PaperarXivGitHubHugging FaceLID, POS, NER, MLI, MT, and TN models trained on this dataset.
Models page →