OpenProblemBench: Benchmarking AI on Open Problems in the Foundational Theoretical Sciences
Quick Answer
OpenProblemBench introduces a benchmark of 82 unresolved scientific problems, assessing AI models' capabilities in foundational theoretical sciences.
Quick Take
GPT-6-Astra achieved a mean solve rate of 14.0%, outperforming other models like Flash, which scored 2.4-3.7%. This benchmark provides a framework for evaluating AI's contributions to unresolved scientific questions.
Key Points
- OpenProblemBench consists of 82 unresolved problems from mathematics and theoretical physics.
- Evaluator models assess submissions for correctness and completeness without reference solutions.
- GPT-6-Astra achieved the highest mean solve rate of 14.0% among evaluated models.
- Performance improvements linked to problem representation and proof completeness.
- The benchmark aims to explore AI's role in foundational theoretical science.
DeepSignal Analysis
What happened
OpenProblemBench has been introduced as a benchmark consisting of 82 unresolved scientific problems in mathematics and theoretical physics. The benchmark aims to evaluate AI models' capabilities in addressing these problems. GPT-6-Astra achieved a mean solve rate of 14.0%, significantly higher than other models like Flash, which scored between 2.4% and 3.7%.
Key evidence
- OpenProblemBench includes 82 unresolved problems sourced from mathematics and theoretical physics literature.
- GPT-6-Astra achieved a mean judged solve rate of 14.0%, outperforming other models such as Flash, which scored 2.4-3.7%.
- The benchmark allows for independent evaluation of AI submissions without reference solutions, assessing correctness and completeness.
Why it matters
This benchmark represents a significant step in evaluating AI's potential contributions to foundational theoretical sciences. By focusing on unresolved problems, it encourages the development of AI systems that can assist in advancing scientific knowledge. The varying performance rates among models highlight the challenges AI faces in this domain and the need for further research and improvement.
Paper Resources
📖 Reader Mode
~2 min readAbstract:The next frontier for artificial general intelligence is tackling unresolved scientific problems, calling for benchmarks that assess progress beyond established knowledge. We introduce OpenProblemBench, a benchmark of 82 unresolved problems drawn from the mathematics and theoretical physics literature. Each problem supplies the research context, assumptions, and prior progress needed to investigate the question. We select problems whose proposed solutions admit comparatively clear checks of their decisive mathematical or computational claims. Four evaluator models independently assess the correctness, completeness, and degree of progress of each submission without reference solutions. Across seven evaluated configurations, GPT-6-Astra achieves the highest mean judged solve rate of 14.0%, compared with 5.5-6.7% for the evaluated full-size open models and 2.4-3.7% for Flash models. Case comparisons connect stronger outcomes to changes in problem representation, general arguments that extend beyond finite evidence, and proofs of the steps needed to complete a solution. By grounding evaluation in questions arising from the research literature, OpenProblemBench provides a setting for investigating the capabilities and limitations of AI as a contributor to foundational theoretical science.
| Comments: | 21 pages, 4 figures, including appendices |
| Subjects: | Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2610.11118 [cs.AI] |
| (or arXiv:2610.11118v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2610.11118 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Sihan Hu [view email]
[v1]
Thu, 8 Oct 2026 02:43:51 UTC (80 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.