Transformer-based segmentation of prosodic boundaries in Brazilian Portuguese
Quick Answer
The paper introduces SAMPA, a Whisper-based segmenter for Brazilian Portuguese that achieves competitive prosodic boundary detection, with F1 scores of 0.731 on a held-out test split and 0.796 on the MuPe-Diversidades dataset.
Quick Take
This model outperforms traditional methods by leveraging deep learning techniques, specifically fine-tuning Whisper large-v3 on the NURC-SP dataset.
Key Points
- SAMPA uses Whisper large-v3, fine-tuned on NURC-SP dataset.
- Achieves F1 scores of 0.731 and 0.796 on different test sets.
- Outperforms traditional rule-based segmentation methods.
- Employs n-gram and acoustic-visual analyses for boundary detection.
- Demonstrates effectiveness in detecting prosodic cues.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Automatic prosodic segmentation identifies boundaries between speech units from acoustic and linguistic evidence. Although recent deep learning approaches have produced strong results for English, automatic segmentation for Brazilian Portuguese (BP) still relies mostly on rule-based or traditional machine-learning methods. This paper presents SAMPA, a Whisper-based segmenter that transcribes BP speech while inserting explicit markers for terminal prosodic boundaries. We fine-tune Whisper large-v3 on manually segmented recordings from the NURC-SP dataset and evaluate different training and test-time filtering configurations, including out-of-distribution testing on the MuPe-Diversidades dataset. SAMPA achieves competitive boundary-detection performance across settings, with the best models reaching F1=0.731 on the held-out test split and F1=0.796 on MuPe-Diversidades. Finally, through n-gram and acoustic-visual analyses, we show that our model follows morphosyntactic, semantic, and prosodic cues for detecting prosodic boundaries.
| Comments: | 6 pages, 5 figures, submitted to an IEEE conference |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.07408 [cs.CL] |
| (or arXiv:2607.07408v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.07408 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Julio Cesar Galdino [view email]
[v1]
Wed, 8 Jul 2026 13:41:43 UTC (886 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.