A Cross-lingual Comparison of Human and Classification Model Entrainment Behavior in Code-switched Speech Settings
Quick Answer
This study analyzes conversational entrainment in code-switched speech across Mandarin-English, Hindi-English, and Spanish-English dialogues, revealing that while lexical entrainment is consistent, acoustic-prosodic features vary contextually.
Quick Take
Classification models, including classical and Transformer-based classifiers, detect entrainment but prioritize different features than humans, highlighting challenges for developing naturalistic conversational agents.
Key Points
- Lexical entrainment is consistent across language pairs in code-switched speech.
- Acoustic-prosodic entrainment shows significant context-specific variation.
- Classical and Transformer classifiers detect entrainment but misprioritize features.
- Study introduces a human-grounded framework for evaluating multilingual models.
- Findings suggest challenges for creating naturalistic code-switched conversational agents.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Conversational entrainment is well-studied in monolingual and written contexts, but remains underexplored in spoken code-switching (CSW). We present a novel cross-lingual analysis of entrainment in Mandarin-English, Hindi-English, and Spanish-English dialogue and show that, while lexical entrainment generalizes across language pairs, entrainment over acoustic-prosodic and CSW style aspects exhibits context-specific variation. We build on these findings by asking whether classification models capture these human behavioral patterns. Applying feature importance and ablation analyses, we find that classical and Transformer-based classifiers detect entrainment reasonably well but consistently prioritize features other than those most salient to human entraining behavior. Our approach introduces a human-grounded framework for evaluating model decision-making in multilingual stylistic contexts, and suggests future challenges for developing conversational agents capable of producing naturalistic code-switched speech.
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.25202 [cs.CL] |
| (or arXiv:2607.25202v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.25202 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Debasmita Bhattacharya [view email]
[v1]
Tue, 28 Jul 2026 02:16:41 UTC (708 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.