What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR
Quick Answer
The study reveals that ASR systems often misjudge atypical speech due to conflating verbatim and intended references.
Quick Take
Benchmarking 11 models, including encoder-decoder and CTC types, shows significant performance disparities, emphasizing the need for appropriate transcription references in evaluations.
Key Points
- ASR systems underperform on atypical speech due to dual transcription references.
- 11 ASR models were benchmarked, revealing significant performance disparities.
- Verbatim references maintain fidelity, while intended references remove disfluencies.
- Model rankings change drastically based on the transcription style used.
- Selecting the right transcription reference is crucial for accurate ASR evaluations.
Paper Resources
📖 Reader Mode
~2 min readAbstract:ASR systems have been often reported to underperform on atypical speech. An often conflated compounding factor is the existence of two valid transcription references: verbatim (actual produced speech, including repetitions/prolongations) and intended (the canonical form of the text with disfluencies removed) in atypical speech recognition depending on context and use-case. Most ASR evaluations conflate this duality into a single ground truth and reward systems that delete disfluencies, ignoring verbatim faithfulness. We benchmark 11 ASR models from encoder-decoder, CTC and transducer families using both verbatim and intended references on atypical stuttered speech as a case study. Our quantitative assessment underlines the disparity in model performance and rankings using the two transcript styles. Through this analysis, we highlight the importance of selecting a suitable transcription reference for valid model selection depending on the use-case, particularly for atypical ASR.
| Comments: | 5 pages, 2 figures, accepted at Interspeech 2026 |
| Subjects: | Computation and Language (cs.CL); Human-Computer Interaction (cs.HC) |
| Cite as: | arXiv:2606.31112 [cs.CL] |
| (or arXiv:2606.31112v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.31112 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hawau Olamide Toyin [view email]
[v1]
Tue, 30 Jun 2026 04:15:39 UTC (161 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.