QuanLing: Cross-Branch Validation of Language Distance Quantification on Western Romance
Quick Answer
This paper shows that The QuanLing framework extends its language distance quantification to Western Romance languages, revealing that Portuguese and Spanish are closest (LaBSE distance 0.0229), while French and Italian are most distant (0.0338).
Quick Take
This study confirms the model's robustness across different language branches and highlights French's higher MLM predictability at 36.12%.
Key Points
- QuanLing combines language distance metrics and property analysis for robust quantification.
- Using 150 parallel sentences, LaBSE distances were computed for four Western Romance languages.
- Portuguese and Spanish had the closest distance, while French and Italian were the most distant.
- The study shows Western Romance languages have a wider absolute distance span than North Germanic.
- French's MLM predictability is significantly higher than Italian's, reflecting orthographic differences.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Quantifying language distance among closely related languages remains a core challenge in quantitative linguistics. Our previous work [1] introduced QuanLing (Quantitative Linguistics via Pretrained Language Models), a quantitative framework combining language distance metrics (sentence embedding distance, tokenization fragmentation rate) with language property analysis (MLM prediction probability), validated on North Germanic (Danish, Norwegian Bokmål, Swedish). This paper extends QuanLing to Western Romance--French, Portuguese, Spanish, Italian--testing cross-branch applicability with the same metric family and aggregation protocol as our North Germanic study, adapted for four languages (English anchor, quadruplet construction). Using 150 four-language parallel sentences, we compute LaBSE sentence embedding distances, tokenization fragmentation rates from four monolingual BERT tokenizers, and mBERT masked language model mutual intelligibility. Results show that Portuguese--Spanish are closest (LaBSE distance 0.0229), French--Italian most distant (0.0338); LaBSE and mBERT rankings agree on 4 of 6 pairs, confirming cross-model robustness. Western Romance shows a wider absolute distance span than North Germanic (0.011 vs. 0.008) but comparable relative ratios (1.48 vs. 1.67), consistent with longer divergence time. French exhibits notably higher MLM predictability (36.12% top-1 accuracy vs. 29.28% for Italian), reflecting its orthography--phonology decoupling. This cross-branch validation provides further evidence for QuanLing's generalizability beyond a single language branch.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08851 [cs.CL] |
| (or arXiv:2610.08851v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08851 arXiv-issued DOI via DataCite |
Submission history
From: Yiping Bai [view email]
[v1]
Sat, 3 Oct 2026 03:11:29 UTC (526 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.