Multilingual & Multimodal Intelligence

Code-mixed NLP

ANRF-supported benchmarks, annotation tools, and models for Hindi–English code-mixing (language identification, POS, NER, translation, and related tasks), February 2023 – February 2026.

Theme
Multilingual & Multimodal Intelligence
Agency
ANRF
Period / amount
Feb 2023 – Feb 2026 · ₹47.674 lakh

Publications

COMI-LINGUA: Expert Annotated Large-Scale Dataset for Multitask NLP in Hindi-English Code-Mixing

Sheth, Beniwal, Singh. Findings of EMNLP 2025.

PaperarXiv

Commentator: A Code-mixed Multilingual Text Annotation Framework

Sheth, Nisar, Prajapati, Beniwal, Singh. EMNLP 2024 System Demonstrations.

PaperGitHub

Datasets

COMI-LINGUA

Expert-annotated Hindi–English code-mixing (181k+ instances) for LID, POS, NER, normalisation, and translation. Findings of EMNLP 2025.

PagePaperarXivGitHubHugging Face

Models

Task models trained on COMI-LINGUA. Weights on Hugging Face.

COMI-LINGUA-LID

Language identification (token classification).

Hugging Face

COMI-LINGUA-MLI

Matrix-language identification.

Hugging Face

Commentator

Code-mixed multilingual text annotation framework. EMNLP 2024 System Demonstrations.

Commentator: A Code-mixed Multilingual Text Annotation Framework

Sheth, Nisar, Prajapati, Beniwal, Singh. Supports language identification, POS tagging, NER, matrix-language identification, translation, and text normalisation.

PagePaperGitHub

All datasets and models →