From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin
Quick Answer
This study adapts NVIDIA's Nemotron 3.5 ASR model for Kikuyu, Dholuo, and Kalenjin, achieving WERs of 42.97% and 33.98% for Kikuyu and Dholuo respectively.
Quick Take
The adaptation addresses challenges like orthographic inconsistency and speaker imbalance, while also reporting negative findings on non-speech labels and over-generation issues.
Key Points
- Kikuyu and Dholuo models achieve 42.97% and 33.98% WER respectively.
- Dholuo records 9.59% CER under historical label policy.
- Kalenjin's adaptation is ongoing, with 68.74% WER on a diagnostic subset.
- The study emphasizes data-centric adaptation without discarding streaming constraints.
- Negative findings include issues with non-speech labels and short-utterance over-generation.
DeepSignal Analysis
What happened
The study adapts NVIDIA's Nemotron 3.5 ASR model for Kikuyu, Dholuo, and Kalenjin languages. The adapted models achieved word error rates (WER) of 42.97% for Kikuyu and 33.98% for Dholuo. Challenges included orthographic inconsistency and speaker imbalance.
Key evidence
- The Kikuyu model achieved a WER of 42.97%, while the Dholuo model reached 33.98%.
- Dholuo recorded a character error rate (CER) of 9.59% and a no-space CER of 8.13%.
- Kalenjin is still under development, with a WER of 68.74% on a specific diagnostic subset.
Why it matters
This adaptation addresses significant challenges in ASR for African languages, such as orthographic inconsistency and speaker imbalance. The findings highlight the complexities involved in developing language-specific systems from multilingual models. The reported negative findings on non-speech labels and over-generation issues also indicate areas needing further research.
What to watch
Future developments in the Kalenjin model are crucial, as it currently shows a high WER. Additionally, monitoring the impact of the reported negative findings on overall model performance will be important. The study's lack of state-of-the-art claims suggests that further validation against public benchmarks is necessary.
Paper Resources
Source Excerpt
Automatic speech recognition (ASR) for African languages is constrained by orthographic inconsistency, annotation artifacts, missing audio, speaker and domain imbalance, and evaluation procedures that differ from deployment. We present an end-to-end engineering study adapting NVIDIA Nemotron 3. 5 ASR Streaming 0. 6B to Kikuyu, Dholuo, and Kalenjin. Starting from a Kenyan Swahili-adapted checkpoint, we retain its cache-aware FastConformer RNN-T, prompt conditioning, and streaming decoder during ful
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →RF-Agent: A Practical Framework for Building Language Agents for RFIC Design
RF-Agent introduces a novel framework for RF circuit design using , creating a unique RF-domain reasoning dataset with over 11,000 samples. The study reveals that domain-specific supervised fine-tuning and semantic retrieval strategies significantly enhance RF reasoning performance, particularly for smaller models.