Dialect-Robust Speech Language Models with Synthetic Pseudo-Dialect Augmentation
Quick Answer
This study presents a novel approach to enhance Speech Language Model (SLM) performance on dialects by synthesizing pseudo-dialect speech using standard-language TTS, achieving improved translation scores for Japanese (from 25.38 to 26.24) and German (from 31.57 to 32.47).
Quick Take
The method also incorporates intermediate standard-text prediction, boosting performance further for Japanese to 28.26 and Chinese from 11.67 to 16.37, demonstrating scalability across languages without requiring dialect-specific speech resources.
Key Points
- Pseudo-dialect speech synthesis requires no real dialect speech data.
- Intermediate standard-text prediction acts as semantic normalization.
- Japanese translation scores improved from 25.38 to 26.24.
- German translation scores increased from 31.57 to 32.47.
- Chinese performance rose from 11.67 to 16.37 with the new approach.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Speech Language Model (SLM) performance often degrades on dialects due to data scarcity. Conventional text-to-speech (TTS) augmentation struggles to cover diverse dialects as it requires a certain amount of real dialect speech. We propose synthesizing pseudo-dialect speech by converting LLM-generated dialect text via a standard-language TTS model, requiring zero real dialect speech. Additionally, we introduce intermediate standard-text prediction during training, acting as semantic normalization for downstream tasks. We evaluate dialect understanding via dialect-to-English speech translation across Japanese, German, and Chinese dialects. Compared to synthetic standard speech baselines, pseudo-dialect augmentation improves scores for Japanese (from 25.38 to 26.24) and German (from 31.57 to 32.47). Furthermore, the intermediate standard-text prediction effectively bridges the semantic gap, boosting performance to 28.26 for Japanese and from 11.67 to 16.37 for Chinese. These results suggest that our approach scales to various languages without requiring speech resources specific to each dialect.
| Comments: | 7 pages, 1 figure, 7 tables. Accepted to IEEE SLT 2026 |
| Subjects: | Computation and Language (cs.CL); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2610.09321 [cs.CL] |
| (or arXiv:2610.09321v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09321 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Shunsuke Mitsumori [view email]
[v1]
Wed, 7 Oct 2026 02:25:43 UTC (190 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.