Diagnosing Correctness Probes under Self-Judgement Confounding
Quick Answer
This study reveals that language models' self-judgement (SJ) often misaligns with objective correctness (OC), leading to incorrect ranking of responses.
Quick Take
In experiments with four instruction-tuned models, SJ consistently outperformed OC in predicting correct outputs, suggesting that SJ-associated polarity is more transferable across tasks, undermining the reliability of OC semantics.
Key Points
- Conflict cases show OC and SJ can predict opposite output orderings.
- SJ direction transfers above chance across four models with up to 14B parameters.
- OC direction consistently underperforms in expected ordering across conditions.
- Transferability of SJ-associated polarity does not confirm objective correctness semantics.
- Findings persist under various controls, including answer likelihood and sequence length.
DeepSignal Analysis
What happened
The study investigates the discrepancies between self-judgement (SJ) and objective correctness (OC) in language models. It finds that SJ often leads to higher rankings of responses than OC, particularly in high-confidence cases. This misalignment raises questions about the reliability of OC as a measure of correctness.
Key evidence
- The research involved four instruction-tuned models with up to 14 billion parameters, examining the relationship between SJ and OC.
- In cases of high-confidence disagreement, responses endorsed by SJ were frequently ranked above those deemed correct by OC.
- The study found that SJ-associated direction consistently transferred across tasks, while OC-associated direction showed below-chance performance in expected ordering.
Why it matters
This research highlights a critical issue in evaluating language model outputs, as reliance on SJ may mislead assessments of correctness. The findings suggest that the current methods for determining OC may not be robust, potentially affecting applications that depend on accurate output evaluation. Understanding this discrepancy is essential for improving model reliability and trustworthiness.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Hidden-state readouts can predict whether language-model outputs are correct, but objective correctness (OC) usually agrees with the model's own self-judgement (SJ), leaving the decoded signal semantically ambiguous. We construct conflict cases in which OC and SJ predict opposite readout orderings. On high-confidence disagreements, conventional correctness-labelled contrasts often rank incorrect/self-endorsed responses above correct/self-rejected responses, following SJ rather than OC. We estimate factorial SJ- and OC-associated directions and evaluate their polarity across mathematical reasoning and factual recall. Across four instruction-tuned models up to 14B parameters, the SJ-associated direction transfers above chance in both cross-domain directions for every model, whereas the OC-associated direction has a below-chance point estimate for the expected OC ordering in every corresponding condition. This transfer asymmetry develops across middle-to-late layers, persists under answer-likelihood, sequence-length, and null-direction controls, and extends to MMLU and binary TruthfulQA without target-domain direction fitting. Across the studied models and diagnostic subsets, the most reliably transferable component preserves SJ-associated polarity. Transferability alone therefore does not establish objective-correctness semantics.
| Subjects: | Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.16799 [cs.CL] |
| (or arXiv:2607.16799v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16799 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yilong Lu [view email]
[v1]
Sat, 18 Jul 2026 12:45:49 UTC (961 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.