Predicting Poets' Origins from Verse: A Computational Analysis of Regional Linguistic Fingerprints in the Complete Tang Poems
Quick Answer
This paper shows that A computational analysis of Tang-dynasty poetry reveals geographic origins through linguistic fingerprints, achieving 0.69 accuracy in predicting poet regions using character n-gram TF-IDF.
Quick Take
The study highlights a distance-decay effect in poetic language and temporal variations in regional separability, suggesting interpretable machine learning can generate hypotheses for literary history.
Key Points
- Model predicts poet origins with 0.69 accuracy, surpassing the 0.53 baseline.
- Linguistic distance correlates with geographic distance (Mantel r=0.40).
- South/North separability varies over time, strongest in Late Tang.
- Misclassifications reflect historical prestige of northern court idiom.
- GuwenBERT transformer matches TF-IDF but does not outperform it.
Paper Resources
📖 Reader Mode
~2 min readAbstract:We ask whether the geographic origin of Tang-dynasty poets leaves a detectable linguistic trace in their work. Aggregating every poem attributed to each author in the Complete Tang Poems (Quan Tang Shi) and linking poets to their administrative circuit of origin via the China Biographical Database (CBDB), we build a poet-level corpus of 357 poets across the ten Tang circuits and frame origin prediction as multi-class classification. Using character $n$-gram TF-IDF together with interpretable domain features (imagery, season, and allusion), classical and neural models predict a poet's broad region (South vs.\ North) at $0.69$ accuracy, well above the $0.53$ majority baseline, and finer circuit-level origin above chance. Beyond classification, three findings emerge. (i) Linguistic distance between circuits grows with geographic distance (Mantel $r=0.40$, $p\approx0.09$ over nine circuits), evidence of a distance-decay effect in poetic language. (ii) The signal interacts with time: South/North separability is at chance in the High Tang and strongest in the Late Tang, consistent with court-driven homogenization at the empire's height followed by regional divergence. (iii) The model's confident errors are historically meaningful -- in the Early Tang, every misclassification is a southern poet read as northern, reflecting the prestige of the northern court idiom. We further show that, when given the whole corpus through a hierarchical frozen-encoder representation, a classical-Chinese transformer (GuwenBERT) only matches -- not beats -- simple TF-IDF, and that combining them adds nothing, indicating that character $n$-grams already capture the regional signal. Our results position interpretable machine learning as a hypothesis generator for literary history.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.24093 [cs.CL] |
| (or arXiv:2606.24093v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2606.24093 arXiv-issued DOI via DataCite |
Submission history
From: Chi-Sheng Chen [view email]
[v1]
Tue, 23 Jun 2026 03:17:44 UTC (522 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.