Auditing Question-Order Effects in Large Language Models with the QQ Equality: Mechanism Characterization and a Saturation Caveat
Quick Answer
This study develops the QQ equality as an audit criterion for large language models (LLMs), revealing that forced-binary next-token log-probabilities are insufficient for distribution-level QQ audits.
Quick Take
The research indicates that 17 out of 18 item pairs were saturated, suggesting a need for pre-specified saturation diagnostics in model evaluations.
Key Points
- Developed QQ equality as an audit criterion for autoregressive .
- Forced-binary next-token log-probabilities inadequate for QQ audits.
- 17 out of 18 item pairs showed saturation in empirical tests.
- Methodology includes audit-logged pipelines and saturation diagnostics.
- Recommends pre-specified diagnostics for next-token distribution evaluations.
DeepSignal Analysis
What happened
The study introduces the QQ equality as a criterion for auditing large language models (LLMs), finding that forced-binary next-token log-probabilities are inadequate for comprehensive QQ audits. In tests, 17 out of 18 item pairs showed saturation, indicating a need for saturation diagnostics in model evaluations.
Key evidence
- The QQ equality serves as an audit criterion for sequential judgments in autoregressive LLMs, revealing limitations in current evaluation methods.
- In empirical tests on an open-weight instruction-tuned model, 17 out of 18 item pairs were found to be saturated, suggesting near-deterministic responses.
- The study recommends implementing pre-specified saturation diagnostics whenever next-token distributions are analyzed as survey-response distributions.
Why it matters
This research highlights significant limitations in the evaluation of LLMs, particularly regarding the reliability of next-token log-probabilities. The findings suggest that existing auditing methods may not adequately capture the complexities of model behavior, which could impact the deployment and trustworthiness of LLMs in practical applications.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Human survey respondents exhibit question-order effects that satisfy the QQ (quantum question) equality, an a priori, parameter-free prediction of the projective quantum question-order model. We develop the QQ equality into an audit criterion for sequential judgments of autoregressive large language models (LLMs). Theoretically, we characterize which mechanism classes satisfy it robustly: marginal-independent kernels satisfy QQ iff all four mismatch transition rates coincide (a class containing the 2D rank-1 projective model with a fixed measurement pair under state variation); a polarity- and position-dependent repetition family is characterized by an exact cross-symmetry condition with closed-form violations; QQ-satisfying behaviors are closed under order-matched mixing; and the rank-2 Contextuality-by-Default criterion translates into audit coordinates as $|\qQQ|\le\OSS$, where $\OSS$ (the order-sensitivity score) totals the order sensitivity of the two marginals. Methodologically, we develop a pre-specified, audit-logged pipeline applicable to any model exposing next-token log-probabilities; it combines worst-case robustness envelopes, sampling-consistency spot checks, full label counterbalancing, and a saturation diagnostic. Empirically, in a first-signal pilot on an open-weight instruction-tuned model under two framings, all pre-specified health gates passed, yet 17/18 and 7/8 item pairs, respectively, were saturated (near-deterministic), and no item was certified residually contextual. Forced-binary next-token log-probabilities were thus inadequate for distribution-level QQ audits under the tested model and prompting conditions; we recommend pre-specified saturation diagnostics whenever next-token distributions are treated as survey-response distributions.
| Subjects: | Computation and Language (cs.CL); Artificial Intelligence (cs.AI); Quantum Physics (quant-ph); Methodology (stat.ME) |
| Cite as: | arXiv:2607.17219 [cs.CL] |
| (or arXiv:2607.17219v1 [cs.CL] for this version) | |
| https://doi.org/10.48550/arXiv.2607.17219 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Pilsung Kang [view email]
[v1]
Sun, 19 Jul 2026 12:22:06 UTC (84 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CL
See more →TriAgent: Divergence-Aware Committees for Cost-Efficient Financial Sentiment Analysis
TriAgent introduces a cost-efficient multi-agent system for financial sentiment analysis, combining VADER, FinBERT, and Qwen2.5. It achieves an F1 score of ~0.87 with significant savings of $9.3M/year at a 10M-user scale compared to GPT-4o-mini, while also detecting hallucinations with an AUC of 0.90.