Error-Aware TF-IDF Retrieval-Augmented Generation for ASR Error Correction
Quick Answer
This study introduces an error-aware TF-IDF retrieval-augmented generation framework for ASR error correction, achieving a hit rate increase from 53.7% to 90.9% on the Persian subset of the FLEURS dataset.
Quick Take
The method reduces the final word error rate from 23.06% to 18.83%, providing significant accuracy improvements with minimal latency.
Key Points
- Proposed a lexical error-aware framework for ASR error correction.
- Increased error-aware hit rate from 53.7% to 90.9% on FLEURS dataset.
- Reduced final word error rate from 23.06% to 18.83%.
- Utilizes a symmetric text normalization module and novel TF-IDF algorithm.
- Addresses phonetic misrecognitions effectively with minimal latency.
Paper Resources
📖 Reader Mode
~2 min readAbstract:End-to-end automatic speech recognition systems frequently hallucinate rare entities and domain-specific terms, especially in low-resource languages. While retrieval-augmented generation frameworks can mitigate these errors using large language models, current architectures face significant challenges. They either rely on standard sparse retrieval that ignores phonetic misrecognitions or utilize heavyweight cross-modal embeddings that introduce high latency. This letter proposes a highly efficient, purely lexical error-aware framework designed to explicitly resolve phonetic and loop hallucinations. Our approach integrates a symmetric text normalization module with a novel error-aware term frequency-inverse document frequency algorithm. By constructing a sparse diagonal penalty matrix based on historical errors, the retriever mathematically prioritizes corrective documents containing specific high-risk misrecognitions. Evaluated on the Persian subset of the FLEURS dataset, our method increased the error-aware hit rate from 53.7% to 90.9%. In end-to-end evaluations, the integrated framework reduced the final word error rate from 23.06% to 18.83%, achieving significant accuracy gains with near-zero inference latency.
| Comments: | 4 pages, 1 figure, 2 tables |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Information Retrieval (cs.IR) |
| Cite as: | arXiv:2606.24915 [cs.CL] |
| (or arXiv:2606.24915v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.24915 arXiv-issued DOI via DataCite |
Submission history
From: Mohammad Aref Jafari-Raddani [view email]
[v1]
Fri, 19 Jun 2026 16:43:31 UTC (32 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.