sk-bench: A Native-First Benchmark for Evaluating Large Language Models in Slovak
Quick Answer
This paper shows that The sk-bench benchmark evaluates large language models for Slovak, revealing that the best open model lags proprietary APIs by 12.6 points.
Quick Take
It introduces 30 datasets and emphasizes the importance of native data, instruction repair, and test-time reasoning for under-resourced languages.
Key Points
- sk-bench includes 30 datasets across ten skill categories for Slovak evaluation.
- The best open model scores 12.6 points lower than proprietary APIs.
- Native data use is crucial where translation fails for under-resourced languages.
- Test-time reasoning boosts scores by 8.5 to 12.5 points for models above 9B.
- Findings suggest planning for instruction repair after language adaptation.
DeepSignal Analysis
What happened
The sk-bench benchmark evaluates large language models specifically for Slovak, highlighting a performance gap between the best open model and proprietary APIs. It introduces 30 datasets tailored for Slovak language evaluation and emphasizes the significance of native data and reasoning capabilities.
Key evidence
- The sk-bench benchmark includes 30 datasets and 33 scored task variants across ten skill categories for evaluating Slovak language models.
- The best open model scored 12.6 points lower than proprietary APIs, indicating a notable performance gap.
- Test-time reasoning improved scores by 8.5 to 12.5 points for models with 9 billion parameters and above.
Why it matters
This benchmark is crucial for advancing the evaluation of language models in under-resourced languages like Slovak, which has been overlooked in multilingual benchmarks. By focusing on native data and specific evaluation metrics, it provides insights that could enhance model performance and applicability in similar linguistic contexts.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAuthors:Marek Šuppa, Ivan Vykopal, Andrej Ridzik, Kristián Sopkovič, Natália Kňažeková, Jaroslav Kopčan, Miroslav Blšták, Viktória Ondrejová, Daniel Hládek, Michal Gregor, Martin Tamajka, Marián Šimko
Abstract:Multilingual LLM benchmarks omit Slovak, a morphologically rich West Slavic language of five million speakers, or cover it only by machine translation. We present sk-bench, a native-first Slovak benchmark with 30 datasets (33 scored task variants) across ten skill categories. Eleven resources are introduced or first packaged for generative-LLM evaluation, including IFEval-SK with Slovak-adapted instruction checkers and native Chiby/SKJ1 resources for Slovak grammar and morphology. We evaluate 55 open- and closed-weights models under one harness. The best open model trails proprietary APIs by 12.6 points. Model rankings are similar for native and translated closed-form data ($\rho\geq0.98$), though translation separates the strongest models less well. By contrast, human-authored and LLM-generated QA questions rank models differently ($\rho=0.72$). For Qwen3-14B, continued Slovak pretraining lowers the overall score by 13.9 points. A small instruction set restores three quarters of that loss. Test-time reasoning improves scores by 8.5 to 12.5 points for models of 9B and above. Together, these findings suggest four design lessons for other under-resourced languages: use native data where translation fails, plan instruction repair after language adaptation, enable test-time reasoning before scaling up, and avoid overinvesting in target-language prompts. We release the data and code at this https URL
| Comments: | Accepted to EMNLP 2026 Main |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| ACM classes: | I.2.7 |
| Cite as: | arXiv:2610.09152 [cs.CL] |
| (or arXiv:2610.09152v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09152 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Marek Šuppa [view email]
[v1]
Tue, 6 Oct 2026 21:51:22 UTC (236 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.