Do Coding Agents Need Executable World Models, Simplification, and Verification to Solve ARC-AGI-3?
Quick Answer
The study evaluates four Codex-based agents on ARC-AGI-3, revealing that while all variants improve with stronger models and reasoning effort, the textual variant outperformed the flexible executable model in certain settings.
Quick Take
The complete verification treatment consistently ranked highest but required more resources, achieving 99% RHAE in follow-ups with gpt-5.6-sol.
Key Points
- Four Codex-based agents tested: textual baseline, executable model, and verification variants.
- Textual variant outperformed flexible executable model in gpt-5.5 settings.
- Complete verification treatment ranked first in all settings, using more resources.
- Simplification improved performance in three out of four model-effort settings.
- gpt-5.6-sol variant solved all public games with 99% RHAE, using fewer actions than humans.
Paper Resources
📖 Reader Mode
~3 min read
Bibliographic and Citation Tools
Code, Data and Media Associated with this Article
Demos
Recommenders and Search Tools
arXivLabs: experimental projects with community collaborators
arXivLabs is a framework that allows collaborators to develop and share new arXiv features directly on our website.
Both individuals and organizations that work with arXivLabs have embraced and accepted our values of openness, community, excellence, and user data privacy. arXiv is committed to these values and only works with partners that adhere to them.
Have an idea for a project that will add value for arXiv's community? Learn more about arXivLabs.
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.