Tiny-Scale Chinese BERT Pretraining: A Controlled Comparison of MLM, WWM, and MacBERT Strategies
Quick Answer
This study compares three pretraining strategies—MLM, WWM, and MacBERT—on a tiny-scale Chinese BERT model (8.7M parameters).
Quick Take
MLM outperforms in overall intrinsic performance, winning 3 out of 5 evaluation dimensions, while WWM shows significant improvements in perplexity and hit rate. MacBERT's performance degrades severely under limited synonym conditions, highlighting the unreliability of training loss as a sole metric.
Key Points
- MLM achieves the best overall performance, winning 3 out of 5 evaluation dimensions.
- WWM shows a 39.5% improvement in perplexity compared to MLM.
- MacBERT's performance degrades significantly under limited synonym conditions.
- Training loss alone is unreliable when mixed replacement strategies are used.
- All models and datasets are publicly available for further research.
DeepSignal Analysis
What happened
A study compared three pretraining strategies—MLM, WWM, and MacBERT—on a tiny-scale Chinese BERT model with 8.7M parameters. MLM outperformed in three out of five evaluation dimensions, while WWM showed notable improvements in perplexity and hit rate. MacBERT's performance declined significantly under limited synonym conditions, indicating potential issues with using training loss as a metric.
Key evidence
- The study evaluated three pretraining strategies on a tiny-scale Chinese BERT model with 8.7M parameters, using a corpus of 1.29M sentences from Chinese Wikipedia.
- MLM achieved the best overall performance by winning 3 out of 5 evaluation dimensions, while WWM excelled in perplexity with a 39.5% improvement.
- MacBERT's performance degraded severely under limited synonym conditions, with a perplexity of 47.23, indicating that training loss is not a reliable metric.
Why it matters
This research highlights the importance of pretraining strategies in language model performance, particularly at smaller scales. The findings challenge established conclusions from larger models, suggesting that different strategies may yield varying results depending on model size and training conditions. Understanding these nuances is crucial for developing effective language models, especially in resource-constrained environments.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Pretraining strategies significantly impact the quality of language models, yet existing comparisons of Masked Language Modeling (MLM), Whole Word Masking (WWM), and MacBERT-style replacement have focused primarily on base-scale models (>=110M parameters). This paper presents a controlled comparison of these three strategies on a tiny-scale Chinese BERT model (4 layers, 256 hidden dimensions, 8.7M parameters). Under identical architecture, corpus (1.29M sentences from Chinese Wikipedia), and hyperparameters, we train three models from scratch and evaluate them across five intrinsic dimensions: perplexity, MLM hit rate, semantic discrimination, grammatical judgment, and contextual sensitivity. At tiny scale, MLM achieves the best overall intrinsic performance (winning 3 of 5 dimensions), while WWM excels in both perplexity (1.27 vs. 2.10, a 39.5% improvement) and MLM hit rate (22% vs. 16%). Notably, MacBERT under a severely limited synonym dictionary (222 entries, 3.3% coverage) exhibits severe perplexity degradation (47.23, 22x higher than MLM), yielding a ranking (MLM > WWM >> MacBERT) that differs markedly from the established base-scale conclusion (MacBERT > WWM > MLM). We further identify a critical evaluation pitfall: MacBERT achieves the lowest training loss (2.17) yet the highest perplexity (47.23), revealing that training loss alone is unreliable under mixed replacement strategies. All models and corpus are publicly available at this https URL.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.08879 [cs.CL] |
| (or arXiv:2610.08879v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.08879 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yiping Bai [view email]
[v1]
Tue, 6 Oct 2026 07:30:05 UTC (10 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.