SPARCLE: SPeaker-aware Aligned Representations via Contrastive Language Embeddings
Quick Answer
SPARCLE introduces a speaker-aware grapheme representation model that significantly enhances text-to-speech (TTS) generation quality, reducing word error rates by 50% in low-resource settings compared to traditional grapheme-based models.
Quick Take
It aligns graphemes with Wav2Vec2 acoustic representations while considering speaker identity, providing a robust alternative to G2P systems.
Key Points
- SPARCLE enhances grapheme modeling by incorporating speaker-specific acoustic variations.
- The model is trained with a contrastive objective for better alignment with acoustic representations.
- It serves as a replacement for G2P systems in downstream TTS tasks.
- Word error rates are halved in extreme low-resource settings compared to standard models.
- SPARCLE demonstrates superior performance over phoneme-based systems at scale.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Recent advances in speech synthesis have shifted from phoneme representations to direct grapheme modeling. While phonemes address the one-to-many mapping between text and acoustics, they rely on grapheme-to-phoneme (G2P) systems that fail to capture speaker-specific acoustic variation. Prior work demonstrates that grapheme-based models outperform phoneme-based systems at scale, but not in low-resource settings.
In this paper, we propose SPARCLE, a speaker-aware grapheme representation model that enriches characters with their precise acoustic realizations. SPARCLE is trained with a contrastive objective to align graphemes with corresponding Wav2Vec2 acoustic representations while conditioned on speaker identity. The resulting model serves as a replacement to G2P systems for downstream text-to-speech (TTS) tasks. We demonstrate that SPARCLE improves generation quality, reducing word error rates by half in extreme low-resource settings compared to standard grapheme-based models.
| Comments: | 5 Pages, 1 Figure, 2 Tables, Interspeech |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Sound (cs.SD); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2607.01238 [cs.CL] |
| (or arXiv:2607.01238v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.01238 arXiv-issued DOI via DataCite |
Submission history
From: Priyam Mazumdar [view email]
[v1]
Fri, 1 May 2026 17:46:18 UTC (205 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.