Same Question, Different Answers: Evaluating LLM Reliability Beyond Accuracy
Quick Answer
This study evaluates the reliability of large language models (LLMs) by examining their responses to paraphrased questions across 13 models and four benchmarks.
Quick Take
Results indicate that while overall accuracy remains stable, instance-level answers fluctuate significantly, with mismatch rates exceeding 23%. A self-paraphrasing strategy shows promise in improving consistency and performance during inference.
Key Points
- Model outputs vary significantly based on prompt wording, with over 23% mismatch rates.
- Instance-level behavior is less stable despite modest overall accuracy changes.
- Correct answers can often be retrieved from at least one paraphrase of a question.
- A self-paraphrasing strategy can enhance performance and recover latent knowledge.
- Standard accuracy metrics may obscure substantial instability in responses.
DeepSignal Analysis
What happened
A study evaluated the reliability of 13 large language models (LLMs) across four benchmarks by analyzing their responses to paraphrased questions. While overall accuracy remained stable, instance-level answers varied significantly, with mismatch rates exceeding 23%. A self-paraphrasing strategy was found to improve consistency and performance during inference.
Key evidence
- The study examined 13 LLMs across four benchmarks to assess their responses to paraphrased questions.
- Mismatch rates for instance-level answers exceeded 23%, indicating significant variability in model responses.
- A self-paraphrasing strategy was shown to enhance consistency and performance during inference.
Why it matters
Understanding the reliability of LLMs is crucial for their deployment in real-world applications. The findings highlight that high accuracy on standard benchmarks does not guarantee consistent performance across different phrasings. This inconsistency could lead to unreliable outputs in critical applications, necessitating improved evaluation methods that focus on response stability rather than just accuracy.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:Large language models (LLMs) often achieve strong accuracy on benchmarks, yet it remains unclear how reliably they apply this knowledge when the same question is phrased in different but equivalent ways. In this work, we study how model answers change under meaning-preserving paraphrases across factual question answering and mathematical reasoning tasks. Across four benchmarks and 13 models, we find that model outputs frequently depend on the exact wording of the prompt. While overall accuracy typically changes only modestly across paraphrases, instance-level behavior is far less stable: for many questions, models alternate between correct and incorrect answers depending on phrasing, with mismatch rates reaching more than 23%. Conditioning on questions that are answered correctly in their original form reveals even larger failures measured by answer flip rates, showing that single-prompt correctness is often a poor indicator of reliability. At the same time, we find that models often produce a correct answer for at least one paraphrase of a question, suggesting that the underlying knowledge is present but inconsistently retrieved. Building on this observation, we show that a simple self-paraphrasing strategy can partially recover this latent knowledge and improve performance at inference time. Together, these findings suggest that standard accuracy metrics can mask substantial instability, and that evaluating consistency across equivalent inputs provides a clearer picture of LLM reliability.
| Subjects: | Artificial Intelligence (cs.AI); Computation and Language (cs.CL); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.22554 [cs.AI] |
| (or arXiv:2607.22554v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22554 arXiv-issued DOI via DataCite |
Submission history
From: Kazem Faghih [view email]
[v1]
Mon, 18 May 2026 16:45:13 UTC (404 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.