KyrgyzLLM-Bench: Benchmarking Kyrgyz Language Understanding
Quick Answer
This paper shows that The KyrgyzLLM-Bench benchmark suite evaluates 26 LLMs in Kyrgyz, revealing performance gaps and translation artifacts.
Quick Take
Notably, few-shot prompting enhances open-source models in reading comprehension, while proprietary models show inconsistent results. All datasets and evaluation tools are publicly released to advance Kyrgyz NLP research.
Key Points
- KyrgyzLLM-Bench includes KyrgyzMMLU and KyrgyzRC datasets.
- Model rankings transfer from English to Kyrgyz for WinoGrande and BoolQ.
- HellaSwag shows a significant English-Kyrgyz performance gap.
- Few-shot prompting improves open-source models but varies for proprietary ones.
- All datasets and evaluation code are publicly available for future research.
DeepSignal Analysis
What happened
The KyrgyzLLM-Bench benchmark suite evaluates 26 large language models (LLMs) in Kyrgyz, highlighting performance gaps and translation artifacts. It includes two natively authored datasets and translated versions of existing benchmarks. Few-shot prompting improves some open-source models, while proprietary models show inconsistent results.
Key evidence
- KyrgyzLLM-Bench evaluates 26 LLMs in Kyrgyz, revealing performance gaps and translation artifacts.
- The benchmark suite includes two natively authored datasets: KyrgyzMMLU and KyrgyzRC.
- Few-shot prompting enhances reading comprehension for several open-source models but behaves inconsistently for proprietary models.
Why it matters
This evaluation addresses the challenge of assessing LLMs in less-resourced languages like Kyrgyz, where native evaluation data is limited. By providing publicly available datasets and evaluation tools, the study aims to advance research in Kyrgyz natural language processing (NLP). The findings on model performance and translation artifacts can inform future developments in multilingual AI applications.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Evaluating large language models (LLMs) across languages remains challenging, as most multilingual benchmarks rely on translated English datasets, often obscuring linguistic and cultural specificity in the target language. This issue is particularly pronounced for less-resourced languages such as Kyrgyz, where reliable natively authored evaluation data are scarce. Building on previously introduced Kyrgyz-language evaluation datasets, this work reports the first systematic and large-scale evaluation of LLMs in Kyrgyz using the KyrgyzLLM-Bench benchmark suite. KyrgyzLLM-Bench comprises two natively authored datasets$-$KyrgyzMMLU and KyrgyzRC$-$together with carefully translated and manually post-edited versions of WinoGrande, HellaSwag, BoolQ, and TruthfulQA. We evaluate 26 open- and closed-source LLMs under zero-shot and few-shot settings, analyzing model performance, cross-lingual transfer, and the impact of translation artifacts on evaluation reliability. Across families and tasks, model rankings transfer broadly from English to Kyrgyz on WinoGrande and BoolQ, and to a lesser extent on MMLU, while HellaSwag exhibits a substantial English-Kyrgyz performance gap consistent with translation-induced plausibility shifts. Few-shot prompting improves several open-source models on reading comprehension but behaves inconsistently for proprietary models on translated tasks. We publicly release all datasets, evaluation code, and per-model results, and integrate the Kyrgyz tasks into a widely used multilingual evaluation framework to support future research on Kyrgyz NLP.
| Comments: | Preprint; manuscript currently under consideration at a journal |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.17173 [cs.CL] |
| (or arXiv:2607.17173v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.17173 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Anton Alekseev [view email]
[v1]
Sun, 19 Jul 2026 10:17:18 UTC (37 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.