Local Prototype Reconstruction for Text-Compatible Speech-to-LLM Bridge Pretraining
Quick Answer
The study introduces Local Prototype Reconstruction (LPR) as a regularizer for speech-to-LLM bridges, enhancing transfer learning in multilingual ASR and speech translation.
Quick Take
LPR significantly improves performance, particularly in translation and low-resource adaptation, by ensuring bridge tokens are reconstructable from embeddings. This method correlates with downstream gains, indicating its potential for reusable speech-to-LLM interfaces.
Key Points
- LPR requires aligned bridge tokens to be reconstructable from LLM token embeddings.
- Improvements observed in multilingual ASR and speech translation tasks.
- Largest performance gains noted in translation and low-resource adaptation.
- Diagnostic measures lexical manifold compatibility, predicting reusability.
- Next-word prediction does not fully capture token-level compatibility.
DeepSignal Analysis
What happened
The study presents Local Prototype Reconstruction (LPR) as a method to enhance the performance of speech-to-LLM bridges. LPR focuses on ensuring that bridge tokens can be reconstructed from LLM embeddings, which is particularly beneficial for multilingual ASR and speech translation tasks. The method shows significant improvements in translation and low-resource adaptation scenarios.
Key evidence
- LPR is introduced as a regularizer that requires bridge tokens to be reconstructable from LLM token embeddings, enhancing the speech-to-LLM interface.
- The study indicates that LPR leads to improved transfer performance, especially in translation and low-resource adaptation tasks.
- An independent diagnostic correlates with downstream gains, suggesting that lexical manifold compatibility is predictive of the reusability of speech-to-LLM bridges.
Why it matters
The introduction of LPR could reshape how speech-to-LLM systems are designed, moving beyond treating the bridge as a mere connector. By focusing on the geometry of the interface, this method may lead to more effective transfer learning across various tasks, particularly in low-resource settings where traditional methods struggle. The potential for reusability of these interfaces could streamline future developments in multilingual ASR and speech translation.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Speech-to-LLM systems often connect a frozen speech encoder to a frozen large language model (LLM) through a small trainable bridge. The bridge is usually treated as plumbing, but it in fact defines the geometry of the speech-to-LLM interface, and the pretraining objective decides whether that interface provides a reusable initialization for downstream tasks. We study a transferable bridge through two complementary properties: global alignment with the text side, and local lexical manifold compatibility, where bridge embeddings remain close to the frozen LLM's input-embedding neighbourhoods. We make this property measurable with a fixed, head-free, timestamp-free diagnostic that applies to any objective, and show that next-word prediction (NWP) and sentence-level contrastive pretraining do not fully capture token-level lexical compatibility. We then introduce Local Prototype Reconstruction (LPR), a lightweight training-only regularizer that requires each aligned bridge token to be reconstructable from a small neighbourhood of frozen LLM token embeddings, with a hard single-prototype anchor as its limiting case. On multilingual ASR and speech translation, LPR improves transfer, with the largest gains on translation and low-resource adaptation. Crucially, our independent diagnostic correlates with downstream gains across objectives, suggesting that lexical manifold compatibility is predictive of reusability for speech-to-LLM bridges.
| Subjects: | Computation and Language (cs.CL); Sound (cs.SD) |
| Cite as: | arXiv:2610.11159 [cs.CL] |
| (or arXiv:2610.11159v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11159 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Xinnian Zhao [view email]
[v1]
Thu, 8 Oct 2026 03:19:49 UTC (202 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →The "10th Juror": Open-Set Standpoint Screening for Bureaucratic Bias Detection
MARS-Gov introduces a framework for detecting bureaucratic bias in Dutch government documents, achieving a new state-of-the-art F1 score of 0.880. This model outperforms existing zero-shot detectors by 20.2 points and reduces unnecessary interventions to just 2.5%. The framework's dynamic '10th juror' adapts to emerging biases, enhancing legal language processing.