
OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness
Quick Answer
OpenAI's GPT-5.6 Sol scores 38.3% on the ARC-AGI-3 benchmark using a custom test harness, surpassing Anthropic's Opus 5 at 30.2%.
Quick Take
However, this performance relies on OpenAI's unique setup, which retains reasoning and compacts context, while the official test yielded only 7.8%. This highlights the impact of testing environments on model evaluations.
Key Points
- GPT-5.6 Sol achieved 38.3% on -3 using OpenAI's custom harness.
- Opus 5 scored 30.2% under the same benchmark conditions.
- Official testing of GPT-5.6 Sol resulted in only 7.8% due to discarded reasoning.
- OpenAI emphasizes the importance of the technical setup in benchmark results.
- The ARC-AGI-3 benchmark aims to measure pure model performance without external aids.
Source Excerpt
OpenAI counters Anthropic's -3 record: GPT-5. 6 Sol scores 38. 3 percent, but only through its own API with retained reasoning and context compaction. In the official test environment, the model managed just 7. 8 percent. Opus 5 hit its 30. 2 percent without such aids.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

