Where to cut, how deep: BPE and Unigram-LM on chemistry SMILES
Quick Answer
This study compares Byte-Pair Encoding (BPE) and Unigram-LM for tokenizing SMILES in chemistry, revealing they generate nearly disjoint vocabularies with a maximum Jaccard overlap of 0.161.
Quick Take
Unigram-LM produces 29-41% more tokens than BPE, indicating a significant difference in segmentation depth across various corpus types and vocabulary sizes.
Key Points
- BPE and Unigram-LM produce nearly disjoint vocabularies in 22 matched conditions.
- Maximum Jaccard overlap between the two methods is only 0.161.
- Unigram-LM segments molecules into 29-41% more tokens than BPE.
- BPE's segmentation is coarser for 80-99% of molecules compared to Unigram-LM.
- The choice of subword algorithm significantly impacts model performance.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Every chemical language model reading SMILES begins with a tokenizer, yet the field has inherited byte-pair encoding (BPE) from natural language with little scrutiny. In natural language, BPE's principal alternative, Unigram-LM, is known to build structurally different vocabularies. Whether that contrast survives in chemistry was open. We report a controlled comparison of BPE and Unigram-LM over a fixed 165-token chemistry base, at the small vocabulary sizes where token embeddings are learnable, across three corpus typologies (diverse, drug-like, natural-products) and both pre-tokenization boundary policies. The two do not converge. In all 22 matched conditions they build near-disjoint subword vocabularies: cross-algorithm Jaccard overlap on the learned pieces never exceeds 0.161, and at most 0.05 once weighted toward the high-frequency pieces a model updates most. Unigram-LM also segments held-out molecules into 29-41% more tokens; the arms largely agree on where to cut but not how deeply, so BPE's segmentation is a strict coarsening of Unigram-LM's on 80-99% of molecules. The separation holds across corpus, boundary, and vocabulary size, persisting even at eight times that scale. The subword algorithm is therefore a modeling decision, not a free default. The study trains no language models.
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG); Biomolecules (q-bio.BM) |
| Cite as: | arXiv:2607.05691 [cs.CL] |
| (or arXiv:2607.05691v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.05691 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hunter Heidenreich [view email]
[v1]
Mon, 6 Jul 2026 23:16:51 UTC (1,009 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.