Evaluating Interactive Reasoning in Large Language Models: A Hierarchical Benchmark with Executable Games
Quick Answer
This paper shows that A new benchmark evaluates interactive reasoning in large language models (LLMs) through 474 executable games, revealing significant performance variances.
Quick Take
The study shows that contextual perturbations reduce success rates moderately, while counterfactual reasoning leads to more substantial declines in performance across various .
Key Points
- The framework evaluates reasoning as active evidence acquisition and belief updating.
- Results indicate large differences in success rates and interaction efficiency among LLMs.
- Contextual perturbations cause moderate declines in performance.
- Counterfactual revision leads to significantly larger drops in success rates.
- The benchmark includes five difficulty levels for comprehensive evaluation.
Paper Resources
📖 Reader Mode
~2 min readAbstract:We introduce a multi-turn interactive framework for reasoning evaluation that treats reasoning as active evidence acquisition and belief updating. Wherein, LLMs receive only the task rules, must issue targeted queries to a hidden environment, integrate partial observations over time, and decide when to submit a final answer. Beyond standard success rate and interaction efficiency, we evaluate contextual robustness under controlled contextual perturbations, and metacognitive adaptation through counterfactual revision and necessity judgment. We instantiate the framework as a benchmark of 474 executable games, each evaluated under five fixed configuration search spaces corresponding to five difficulty levels, and evaluate a broad set of frontier LLMs. Results show that the benchmark is highly discriminative, exposing large differences not only in success rate but also in interaction efficiency. Moreover, we empirically show that contextual perturbations cause moderate but consistent declines, whereas counterfactual revision and necessity judgment lead to much larger drops.
| Comments: | preprint version, under review |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2606.00103 [cs.AI] |
| (or arXiv:2606.00103v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2606.00103 arXiv-issued DOI via DataCite |
Submission history
From: Mingyuan Fan [view email]
[v1]
Tue, 26 May 2026 09:12:30 UTC (34 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.