Multilingual Steering by Design: Multilingual Sparse Autoencoders and Principled Layer Selection
Quick Answer
The study introduces multilingual sparse autoencoders (SAEs) that enhance cross-lingual representations and improve language control in large language models like LLaMA-3.1-8B and Gemma-2-9B.
Quick Take
By employing a principled layer-selection rule, the method stabilizes the balance between language identification accuracy and generation quality, as evidenced by improved SpBLEU and ROUGE-L scores in machine translation and cross-lingual summarization tasks.
Key Points
- Multilingual SAEs trained on diverse data improve language control reliability.
- A new layer-selection rule predicts effective intervention depths without exhaustive searches.
- Evaluated on LLaMA-3.1-8B and Gemma-2-9B with positive results in SpBLEU and ROUGE-L.
- Stabilizes trade-off between language identification accuracy and generation quality.
- Applicable to machine translation and cross-lingual summarization tasks.
Paper Resources
Article Content
From source RSS / original summaryarXiv:2605. 23036v1 Announce Type: new Abstract: Sparse autoencoders (SAEs) enable feature-level mechanistic interpretability and activation steering in (LLMs), but SAE-based language control remains unreliable in multilingual settings: most SAEs are trained on English-only data, and steering layers are chosen heuristically. We address these limitations by advancing a principled, mechanistic account of multilingual language steering with SAEs.
First, we show that training SAEs on multilingual data consistently strengthens cross-lingual representations and yields more reliable, quality-preserving language control across layers and model families. Second, we introduce an \emph{a priori} steering layer-selection rule based on the intersection of multilingual alignment and language separability, which predicts effective intervention depths without exhaustive layerwise search. We evaluate our approach on LLaMA-3.
1-8B and Gemma-2-9B across machine translation and cross-lingual summarization (CrossSumm), using SpBLEU, ROUGE-L, COMET, and LaSE. Our results show that multilingual SAEs combined with intersection-selected layers stabilize the trade-off between language identification accuracy and generation quality, providing a principled, predictive, representation-level account of multilingual SAE steering.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.