Same Text, Different Prediction: Serving-Context Nondeterminism in Text Classifiers
Quick Answer
This study reveals significant nondeterminism in text classifiers, showing that changing batch shape can shift predicted probabilities by up to 56.7 points under bf16 precision.
Quick Take
It highlights that fully generative classifiers exhibit more label changes than discriminative ones, emphasizing the need for fixed serving conditions to ensure reproducibility in text classification.
Key Points
- 180 models were trained across discriminative, pseudo-generative, and fully generative classifiers.
- Label stability does not guarantee score stability; batch shape changes affect predictions significantly.
- Under bf16 precision, predicted probability mass can shift by up to 56.7 percentage points.
- Fully generative classifiers change more labels than discriminative classifiers under similar conditions.
- The study provides conditions and mitigations for achieving label stability in text classification.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Deterministic inference is essential for reliable and trustworthy machine learning. Prior studies of text generation have shown that changing factors such as batch size, batch composition, hardware, or inference engine can alter the generated text, even when the prompt, model parameters, and sampling randomness are fixed. These differences have been attributed in part to floating-point non-associativity, shape-dependent kernel selection, and other implementation-level differences in numerical execution. However, it remains unclear whether, when, and to what extent the same factors affect text classification. We present a systematic study of serving-context non-invariance in text classifiers, which prior work has measured only through generated text. We train 180 models spanning discriminative, pseudo-generative, and fully generative classifier formulations and evaluate each across four categories of serving contexts, holding the checkpoint and the text fixed. Label stability does not imply score stability. Changing only the batch shape changes no labels across fp32 comparisons, yet under bf16 it moves up to 56.7 percentage points of predicted probability mass, with label changes concentrated at small margins. Fully generative classifiers change more labels than their discriminative counterparts under the same serving changes. We derive sufficient conditions for label stability under each serving change and give a separate mitigation for each mechanism. Our results identify and quantify the serving conditions that must be fixed for reproducible text classification.
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2610.09111 [cs.CL] |
| (or arXiv:2610.09111v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2610.09111 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Siva Rajesh Kasa [view email]
[v1]
Tue, 6 Oct 2026 21:02:25 UTC (283 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.