Tokka-Bench: Evaluating Tokenizers Across 100 Natural and 20 Programming Languages
Quick Answer
Tokka-Bench is an open-source framework that evaluates tokenizers across 100 natural and 20 programming languages using five metrics.
Quick Take
It reveals that vocabulary allocation strategies are more crucial than size, and recent programming tokenizers show efficiency convergence despite diverse natural language profiles.
Key Points
- Evaluates seven BPE tokenizers including GPT-2 and Llama 3.1.
- Uses five metrics: bytes per token, unique token coverage, and more.
- Framework and data are publicly available for further research.
- Programming language tokenizers have converged in efficiency.
- Vocabulary allocation strategy impacts tokenizer performance significantly.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models rely on subword tokenizers whose quality varies across languages, yet no standardized multi-metric framework exists for broad comparative evaluation. We introduce Tokka-Bench, an open-source framework that evaluates tokenizers on five complementary metrics -- bytes per token, unique token coverage, subword fertility, word-split rate, and vocabulary composition -- across 100 natural languages (30+ scripts) and 20 programming languages, using language-aware segmentation adapted to each writing system. Comparing seven BPE tokenizers (GPT-2, GPT-4, gpt-oss, Llama 3.1, Gemma 3, Qwen3, and Kimi K2) within individual languages, we find that vocabulary allocation strategy matters more than raw vocabulary size, and that programming-language efficiency has converged among recent tokenizers despite divergent natural-language profiles. The framework, data, and interactive dashboard are publicly available.
| Comments: | 5 pages, 5 figures. Code and data: this https URL. Interactive dashboard: this https URL |
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08794 [cs.CL] |
| (or arXiv:2610.08794v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08794 arXiv-issued DOI via DataCite |
Submission history
From: Ben Gubler [view email]
[v1]
Wed, 25 Mar 2026 21:01:32 UTC (47 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.