Voice Memory for Agentic Speech Recognition
Quick Answer
Voice Memory introduces an inference-only scheme for agentic speech recognition, reducing generative error correction over-corrections from 64% to 35%.
Quick Take
It lowers the weighted word error rate from 8.36% to 7.52% across ten HyPoradise domains, significantly improving performance in areas like air-travel commands. The system operates without changing weights and adds no parameters to the inference path.
Key Points
- Voice Memory reduces generative error correction over-corrections from 64% to 35%.
- Weighted word error rate drops from 8.36% to 7.52% across ten domains.
- Performance gains are notable in air-travel commands, improving from 8.40% to 3.40%.
- The system maintains auditability and portability without changing model weights.
- Demo and example code are available for further research.
Paper Resources
📖 Reader Mode
~2 min readAbstract:We present Voice Memory, a inference-only scheme for agentic speech recognition: at stream time, a frozen corrector reads a single per-domain this http URL and decides per utterance whether to act on the hypothesis or abstain and keep the 1-best. Asynchronously, a score-gated optimizer revises that file through bounded edits, accepting an edit only when it strictly improves a held-out score. Extended from classical ASR-LM framework, we refer this split the listener-thinker architecture; the two roles are coupled only through the memory, so no weights change and the learned skill stays auditable and portable. Restraint turns out to be the operative skill this loop discovers: unconstrained generative error correction (GER) over-corrects, breaking correct tokens on up to 64% of its edits on financial news, and Voice Memory, reduces this rate to 35%. Across ten HyPoradise domains with an open corrector, Voice Memory, lowers weighted word error rate from 8.36% to 7.52% (7.47% with three added in-context examples) without regressing any dataset below its 1-best baseline; gains concentrate where recoverable headroom is largest, including air-travel commands (8.40% to 3.40%) and noisy far-field speech (CHiME-4, 12.69% to 10.46%). The memory transfers across corrector families and adds zero parameters to the inference path. A demo and example code are provided for future studies.
| Comments: | Preprint. Technical report and open source: this https URL |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2607.26410 [cs.CL] |
| (or arXiv:2607.26410v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.26410 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Huck Yang [view email]
[v1]
Wed, 29 Jul 2026 02:42:56 UTC (232 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.