Layer-wise Probing of wav2vec 2.0 and Whisper for Consonant Cluster Reduction in African American English
Quick Answer
This study probes wav2vec 2.0 and Whisper models to analyze consonant cluster reduction (CCR) in African American English (AAE).
Quick Take
Both models accurately differentiate between reduced and canonical forms, revealing that CCR is represented as structured phonological variation rather than mere deletion, impacting automatic speech recognition (ASR) performance.
Key Points
- Layer-wise probing of wav2vec2-base and Whisper-small reveals insights into AAE phonology.
- Both models achieve high accuracy in detecting segmental reduction and restoration tasks.
- Reduced segments maintain cues to underlying stops, indicating structured phonological encoding.
- Findings highlight the need for improved ASR systems for African American English.
- Study contributes to understanding linguistic representation in modern speech models.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Self-supervised and supervised speech models are increasingly used to investigate which linguistic information their internal representations encode, and at what level of abstraction they encode it. One underexplored phenomenon is consonant cluster reduction (CCR) in African American English (AAE), a widespread phonological process and a source of automatic speech recognition (ASR) disparity. To examine how CCR is represented, we conduct speaker-independent layer-wise probing of wav2vec2-base and Whisper-small using two tasks: segmental reduction detection and segmental restoration of underlying cluster identity. Both models distinguish reduced and canonical forms with high accuracy. Crucially, reduced segments retain cues to their underlying stops, indicating that CCR is encoded as structured gradient phonological variation rather than simple segmental deletion. These results demonstrate structured phonological encoding of AAE CCR patterns in modern speech models.
| Comments: | This paper has been accepted for presentation at Interspeech 2026 |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2606.23948 [cs.CL] |
| (or arXiv:2606.23948v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.23948 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hamid Mojarad [view email]
[v1]
Mon, 22 Jun 2026 21:19:03 UTC (163 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.