Evaluation of forced alignment of code-mixed speech: the case of Hindi-English
Quick Answer
The study evaluates the Montreal Forced Aligner's performance on Hindi-English code-mixed speech, achieving a mean error of 4.15ms—ten times lower than monolingual alternatives.
Quick Take
Key improvements stem from bootstrapping strategies and tailored lexicon design, addressing free variation and phonemic boundary detection.
Key Points
- Montreal Forced Aligner shows significant improvement in code-mixed speech alignment.
- Mean error of 4.15ms is ten times lower than monolingual Hindi and English.
- Bootstrapping strategies outperform unmodified lexicons for better accuracy.
- Challenges include orthographic errors and speaker variation in code-mixed speech.
- Effective alignment requires principled lexicon design and code-mixed training data.
Paper Resources
📖 Reader Mode
~1 min readAbstract:Code-mixed speech poses unique challenges to forced alignment: expanded inventories, orthographic errors, and speaker variation. We evaluate forced alignment of Hindi-English code-mixed speech using the Montreal Forced Aligner. We address 2 problems: (1) free variation involving native vs non-native pairs and (2) phonemic boundary detection for mid-utterance English words. Bootstrapping strategies substantially outperform unmodified lexicons. Acoustic models trained on sentence-level code-mixed data achieve a mean error of 4.15ms, ie. ten times lower than monolingual Hindi (38.18ms) or isolated English (37.58ms) alternatives. Principled lexicon design and code-mixed training data are both essential for reliable alignment of bilingual speech.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.25581 [cs.CL] |
| (or arXiv:2607.25581v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.25581 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pamir Gogoi [view email]
[v1]
Tue, 28 Jul 2026 11:08:05 UTC (576 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.