BHARATI: Morphology-Aware Tokenizers for Classical Indian Languages with Subword Fertility Analysis
Quick Answer
BHARATI introduces morphology-aware tokenizers for classical Indian languages, outperforming GPT-2 and multilingual SentencePiece.
Quick Take
The v3 tokenizer reduces sequence length by 90% compared to GPT-2 and 25% against mBART-50, enhancing context for downstream models.
Key Points
- BHARATI tokenizers are trained on a 781 MB corpus across seven languages.
- Subword fertility analysis shows v3 averages 2.6 tokens per IKS term.
- v3 significantly reduces sequence length, improving effective context for models.
- Open licenses for tokenizer models, training scripts, and benchmarks are provided.
- Three tokenizer versions were developed, expanding language support progressively.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages. Sanskrit, Tamil, and other classical Indic languages exhibit agglutinative morphology, productive sandhi (phonological fusion at word boundaries), and domain-specific vocabularies absent from general-purpose training data. This paper presents BHARATI, a set of SentencePiece BPE tokenizers trained on a balanced 781 MB corpus spanning seven languages (English, Hindi, Sanskrit, Tamil, Telugu, Kannada, and Malayalam) with native script support for all languages. We describe three successive tokenizer versions: v1 (English and Sanskrit only, with broken byte-fallback for Tamil), v2 (four-language support with byte-level fallback for southern languages), and v3 (full seven-language native subword coverage). Subword fertility analysis demonstrates that v3 averages 2.6 tokens per Indian Knowledge System (IKS) technical term, compared to 5.25 tokens per term with GPT-2's tokenizer and 3.75 tokens with the multilingual SentencePiece baseline, with the largest gains on a set of reserved IKS terms that are represented as single tokens by construction. On a held-out test set of 490 IKS-domain sentences (70 per language across seven languages, released with the measurement script), v3 reduces sequence length by roughly 90% relative to GPT-2 and byte-level encoding (which lack native Indic subwords) and by approximately 25% relative to the mBART-50 multilingual baseline, averaged across the six Indic languages, directly translating to increased effective context length for downstream language models. The tokenizer models (32,000 vocabulary), training scripts, and evaluation benchmarks are released under open licenses.
| Comments: | 33 pages, 6 figures |
| Subjects: | Computation and Language (cs.CL); Computers and Society (cs.CY); Emerging Technologies (cs.ET); Machine Learning (cs.LG) |
| MSC classes: | 68T50 |
| ACM classes: | I.2.7 |
| Cite as: | arXiv:2607.23319 [cs.CL] |
| (or arXiv:2607.23319v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.23319 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Santhosh Sivasubramani Prof [view email]
[v1]
Sat, 25 Jul 2026 18:23:06 UTC (104 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.