COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing
Sheth, Beniwal, Singh. Findings of EMNLP 2025.
PaperarXivMultilingual & Multimodal Intelligence
ANRF-supported benchmarks, annotation tools, and models for Hindi–English code-mixing (language identification, POS, NER, translation, and related tasks), February 2023 – February 2026.
Sheth, Beniwal, Singh. Findings of EMNLP 2025.
PaperarXivSheth, Nisar, Prajapati, Beniwal, Singh. EMNLP 2024 System Demonstrations.
PaperGitHubExpert-annotated Hindi–English code-mixing (181k+ instances) for LID, POS, NER, normalisation, and translation. Findings of EMNLP 2025.
PagePaperarXivGitHubHugging FaceTask models trained on COMI-LINGUA. Weights on Hugging Face.
Language identification (token classification).
Hugging FacePOS tagging.
Hugging FaceNamed-entity recognition.
Hugging FaceMatrix-language identification.
Hugging FaceCode-mixed translation.
Hugging FaceText normalisation.
Hugging FaceCode-mixed multilingual text annotation framework. EMNLP 2024 System Demonstrations.
Sheth, Nisar, Prajapati, Beniwal, Singh. Supports language identification, POS tagging, NER, matrix-language identification, translation, and text normalisation.
PagePaperGitHub