From a Multilingual Streaming ASR Backbone to Kenyan-Language Systems: Data-Centric Adaptation of Nemotron 3.5 for Kikuyu, Dholuo, and Kalenjin
Quick Answer
This study adapts NVIDIA's Nemotron 3.5 ASR model for Kikuyu, Dholuo, and Kalenjin, achieving WERs of 42.97% and 33.98% for Kikuyu and Dholuo respectively.
Quick Take
The adaptation addresses challenges like orthographic inconsistency and speaker imbalance, while also reporting negative findings on non-speech labels and over-generation issues.
Key Points
- Kikuyu and Dholuo models achieve 42.97% and 33.98% WER respectively.
- Dholuo records 9.59% CER under historical label policy.
- Kalenjin's adaptation is ongoing, with 68.74% WER on a diagnostic subset.
- The study emphasizes data-centric adaptation without discarding streaming constraints.
- Negative findings include issues with non-speech labels and short-utterance over-generation.
DeepSignal Analysis
What happened
The study adapts NVIDIA's Nemotron 3.5 ASR model for Kikuyu, Dholuo, and Kalenjin languages. The adapted models achieved word error rates (WER) of 42.97% for Kikuyu and 33.98% for Dholuo. Challenges included orthographic inconsistency and speaker imbalance.
Key evidence
- The Kikuyu model achieved a WER of 42.97%, while the Dholuo model reached 33.98%.
- Dholuo recorded a character error rate (CER) of 9.59% and a no-space CER of 8.13%.
- Kalenjin is still under development, with a WER of 68.74% on a specific diagnostic subset.
Why it matters
This adaptation addresses significant challenges in ASR for African languages, such as orthographic inconsistency and speaker imbalance. The findings highlight the complexities involved in developing language-specific systems from multilingual models. The reported negative findings on non-speech labels and over-generation issues also indicate areas needing further research.
What to watch
Future developments in the Kalenjin model are crucial, as it currently shows a high WER. Additionally, monitoring the impact of the reported negative findings on overall model performance will be important. The study's lack of state-of-the-art claims suggests that further validation against public benchmarks is necessary.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Automatic speech recognition (ASR) for African languages is constrained by orthographic inconsistency, annotation artifacts, missing audio, speaker and domain imbalance, and evaluation procedures that differ from deployment. We present an end-to-end engineering study adapting NVIDIA Nemotron 3.5 ASR Streaming 0.6B to Kikuyu, Dholuo, and Kalenjin. Starting from a Kenyan Swahili-adapted checkpoint, we retain its cache-aware FastConformer RNN-T, prompt conditioning, and streaming decoder during full-parameter fine-tuning. The study covers corpus auditing, Unicode normalization, split checks, duration filtering, low-rate continuation, validation-based checkpoint selection, true-streaming evaluation, artifact preservation, and isolated serving.
On internal, adaptively consulted evaluation sets excluded from gradient updates at context [56,13], selected Kikuyu and Dholuo models achieve 42.97% and 33.98% WER, respectively. Dholuo records 9.59% CER and 8.13% no-space CER under its frozen historical label policy; Kikuyu records 7.79% no-space CER. Kalenjin remains a work in progress: v1-v reaches 68.74% WER on a 2,411-row clean-v3 diagnostic subset excluding long-pause annotations, digit-bearing references, and targets shorter than three tokens. Its checkpoint selection used a mixed-source validation manifest containing test-origin rows, so the score is not an independent generalization estimate. We also report negative findings involving non-speech labels, short-utterance over-generation, boundary-sensitive WER, and cloud job-lifecycle failures. We make no state-of-the-art claim because the internal sets, repeated consultation, and normalization differ from public benchmarks. This work provides an auditable account of adapting a multilingual streaming model into language-specific systems without discarding streaming constraints.
| Comments: | 56 pages, 2 figures. Extended appendices on corpus construction, streaming evaluation, reproducibility, and deployment |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.18912 [cs.CL] |
| (or arXiv:2607.18912v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.18912 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mark Gatere [view email]
[v1]
Tue, 21 Jul 2026 09:52:57 UTC (54 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.