Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems
Quick Answer
The study introduces a percentile-based evaluation protocol for speech-to-speech AI agents, using over 4000 hours of conversation data to assess prosody and rhythm.
Quick Take
This method improves the calibration of evaluation metrics like $F_0$ expressivity and speech rate, yielding more interpretable results compared to pooled human statistics.
Key Points
- Developed matched reference regimes for key speech metrics including $F_0$ and speech rate.
- Utilized over 4000 hours of dyadic English conversation from the Seamless Interaction dataset.
- Percentile-based evaluation protocol yields flag rates closer to the nominal 10%.
- Matched references provide interpretable deviation directions for speech outputs.
- Complements existing perceptual and user-centered evaluation methods.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Speech-to-speech (S2S) AI agents are advancing rapidly, yet evaluation lacks interpretable speech-native measures for conversational prosody and rhythm. Because $F_0$, speaking rate, articulation rate, and pausing shift with model-predicted speaker traits and interaction state, pooled human statistics can be poorly calibrated for evaluating a particular output. Using 4000+ hours of dyadic English conversation from the Seamless Interaction dataset, we construct matched reference regimes for $F_0$ mean, $F_0$ expressivity, speech rate, articulation rate, pause ratio, and mean pause duration. We then define a percentile-based evaluation protocol: extract the same metrics from an S2S output waveform, compare them to the closest matched human reference stratum, and report percentile deviations or 5th-95th percentile out-of-regime flags. On held-out human rows, pooled references over-flag state-conditioned $F_0$ expressivity and rhythm, while matched references return flag rates closer to the nominal 10% and make deviation direction interpretable. These outputs serve as behavioral plausibility checks that complement, rather than replace, perceptual and user-centered evaluation.
| Subjects: | Computation and Language (cs.CL); Sound (cs.SD); Audio and Speech Processing (eess.AS) |
| Cite as: | arXiv:2606.31055 [cs.CL] |
| (or arXiv:2606.31055v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.31055 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Ashish Hallur [view email]
[v1]
Tue, 30 Jun 2026 02:46:16 UTC (159 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.