Code-Switching Reveals Language Anchoring in Multilingual LLMs
Quick Answer
This paper shows that Multilingual Large Language Models (MLLMs) struggle with Code-Switched (CS) inputs, showing performance degradation due to Anchor Bias.
Quick Take
The proposed CANVAS intervention improves Question Answering (QA) F1 scores across various MLLMs by aligning target-language hidden states with source anchors during inference.
Key Points
- Anchor Bias quantifies language anchoring in MLLMs, revealing a grammar-frame effect.
- Source-framed CS maintains source anchoring, while target-framed CS shows greater QA degradation.
- CANVAS intervention effectively recovers QA F1 scores across diverse MLLMs and CS conditions.
- Internal anchoring signals can mitigate CS inference failures in multilingual models.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Multilingual Large Language Models (MLLMs) are increasingly expected to handle Code-Switched (CS) inputs, yet mixing languages frequently degrades performance relative to source- or target-language monolingual counterparts. To understand this degradation, we use grammar-forced CS as a controlled diagnostic setting for locating CS representations relative to their source and target counterparts. We introduce Anchor Bias, a geometric measure that quantifies language anchoring, whether a CS hidden state aligns closer to its source or target language counterpart. Across diverse MLLMs, Anchor Bias reveals a consistent grammar-frame effect: source-framed CS stays source-anchored, whereas target-framed CS shifts target-ward and shows larger Question Answering (QA) degradation. Motivated by this representational pattern, we propose CANVAS (Contextual Anchor-based Neural Vector Alignment Steering), an inference-time intervention that extracts a source-side canvas from the input and softly steers target-language hidden states toward the source anchor during prefill. CANVAS consistently recovers QA F1 across MLLMs and CS conditions, showing that internal anchoring signals provide an actionable target for mitigating CS inference failures.
| Comments: | 36 pages, 13 figures, 27 tables |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2606.19668 [cs.CL] |
| (or arXiv:2606.19668v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.19668 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Jeonghyun Park [view email]
[v1]
Thu, 18 Jun 2026 00:42:55 UTC (1,256 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.