Evaluation design conditions the expert-vs-auto MeSH gap: a controlled comparison of bag-of-words and BiomedBERT on the Cohen benchmark
Quick Answer
This study compares bag-of-words and BiomedBERT classifiers using the Cohen benchmark, revealing that evaluation design significantly impacts the expert-vs-auto MeSH gap.
Quick Take
The bag-of-words model shows a gap of +0.096 WSS@95% under canonical evaluation, which narrows to +0.021 with 10-fold cross-validation, while BiomedBERT's performance is comparable at +0.020. The findings suggest that classifier performance can vary considerably based on evaluation methods.
Key Points
- Bag-of-words classifier shows a +0.096 WSS@95% gap in canonical evaluation.
- 10-fold cross-validation reduces the bag-of-words gap to +0.021.
- BiomedBERT's performance is comparable at +0.020 WSS@95%.
- Representation asymmetry affects results with 15.1% of inputs exceeding BiomedBERT's token limit.
- Evaluation design significantly alters conclusions about feature sources.
Paper Resources
📖 Reader Mode
~2 min readAbstract:A systematic review begins with someone reading thousands of abstracts to identify the few that are relevant, and classifiers are used to prioritise that reading. Their inputs are often augmented with Medical Subject Headings (MeSH), assigned either by expert indexers weeks or months after publication or by automatic tools at once. To our knowledge the two have not been compared directly as classifier features, and no previous work has asked whether that comparison's outcome depends on how the classifier is evaluated. Using the Cohen et al. (2006) drug-class benchmark on three topics, we characterise a bag-of-words logistic regression classifier (seven reruns) and BiomedBERT (five seeds), then examine how the Statins result changes under alternative designs. Under the canonical 5-fold full-corpus design, the bag-of-words expert-vs-auto gap on Statins is +0.096 WSS@95%. Matching the corpus size to the smaller topics (n = 803) reduces it to +0.033 (95% bootstrap CI includes zero), and 10-fold cross-validation at full size to +0.021 (CI narrowly excludes zero). Under canonical evaluation BiomedBERT gives +0.020, within sampling noise of the bag-of-words 10-fold result. A power analysis indicates a Statins-sized effect would not have been detectable at the Opioids or ADHD variance, so those nulls are design-limited rather than informative. A representation asymmetry remains: 15.1% of Statins inputs exceed BiomedBERT's 512-token limit when expert MeSH terms are appended, so truncation may contribute to the smaller transformer gap, although this cannot be separated from training volume here. In screening pipelines using transformers or 10-fold bag-of-words, the gap on the topics tested is about 0.02 WSS@95%, with CIs spanning zero on at least one bound. More broadly, benchmark conclusions about feature sources can change substantially under reasonable changes to the evaluation design.
| Comments: | 15 pages, 2 figures, 10 tables |
| Subjects: | Computation and Language (cs.CL) |
| Cite as: | arXiv:2607.21685 [cs.CL] |
| (or arXiv:2607.21685v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.21685 arXiv-issued DOI via DataCite |
Submission history
From: Samuel Okoe-Mensah [view email]
[v1]
Thu, 23 Jul 2026 14:37:23 UTC (74 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.