Estimating Uncertainty in Classifier Performance with Applications to Large Language Models and Nested Data
Quick Answer
This paper evaluates confidence interval methods for classifier performance metrics in text classification, highlighting that traditional methods like the Wald interval are often inaccurate.
Quick Take
It proposes improved techniques such as Agresti-Coull and a novel pseudo-count regularized bootstrap, particularly for small datasets and nested data scenarios, enhancing transparency in machine learning applications.
Key Points
- Default methods like Wald interval often yield inaccurate confidence intervals.
- Agresti-Coull and Wilson methods improve accuracy for small sample sizes.
- A novel pseudo-count regularized bootstrap is effective for F1 score calculations.
- Hierarchical bootstrap outperforms cluster bootstrap in moderate text production scenarios.
- Improved interval estimation promotes better validation practices in machine learning.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Researchers increasingly use text classification--supervised models or large language models--to measure constructs from natural language, providing metrics such as recall and precision as evidence of their validity. Yet, though these metrics are point estimates subject to sampling variation, measures of uncertainty are inconsistently reported alongside them. Further, when they are reported, they are often estimated with methods that are not appropriate when relevant labelled datasets are small or performance is high. To increase and improve confidence interval reporting in the field, this paper evaluates confidence interval methods for performance metrics under conditions typical of social science text classification: small to moderate sample sizes, infrequent constructs, and texts nested within individuals. Across simulations, default methods such as the Wald interval and the basic percentile bootstrap are the least accurate, with coverage sometimes far below the nominal 95% level. Accuracy is improved with the use of Agresti-Coull, Wilson, Clopper-Pearson, and a novel pseudo-count regularized bootstrap (which is particularly relevant to the calculation of F1). When texts are nested within individuals, we demonstrate that adjustment for both effective N and the appropriate degrees of freedom is necessary for producing accurate analytic intervals. Among bootstrap intervals, the hierarchical bootstrap is more accurate than the cluster bootstrap when individuals produce a moderate number of texts but overly conservative when individuals produce only a few. By providing guidance to the field on appropriate interval estimation, we aim to improve the transparency of machine learning applications, and to encourage greater attention to the validation sample size at the design stage.
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.26422 [cs.AI] |
| (or arXiv:2606.26422v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.26422 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Kylie Anglin [view email]
[v1]
Wed, 24 Jun 2026 22:24:25 UTC (7,983 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.