
OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 but only with its own custom test harness
Quick Answer
OpenAI's GPT-5.6 Sol scores 38.3% on the ARC-AGI-3 benchmark using a custom test harness, surpassing Anthropic's Opus 5 at 30.2%.
Quick Take
However, this performance relies on OpenAI's unique setup, which retains reasoning and compacts context, while the official test yielded only 7.8%. This highlights the impact of testing environments on model evaluations.
Key Points
- GPT-5.6 Sol achieved 38.3% on -3 using OpenAI's custom harness.
- Opus 5 scored 30.2% under the same benchmark conditions.
- Official testing of GPT-5.6 Sol resulted in only 7.8% due to discarded reasoning.
- OpenAI emphasizes the importance of the technical setup in benchmark results.
- The ARC-AGI-3 benchmark aims to measure pure model performance without external aids.
📖 Reader Mode
~1 min readOpenAI says it can keep up on ARC-AGI-3. After Anthropic's Claude Opus 5 quadrupled the record score on the logic benchmark, OpenAI is now showing that GPT-5.6 Sol hits 38.3 percent with two API settings, beating Opus 5's 30.2 percent.

OpenAI isn't using the official test environment, though. Instead, it runs GPT-5.6 Sol through its own Responses API with "Retained Reasoning," which keeps the model's chain of thought between steps, and "Compaction," which summarizes old context instead of truncating it. In the official harness, GPT-5.6 Sol scored just 7.8 percent because the model's reasoning gets discarded after each action.
OpenAI argues that benchmarks never measure just the model but also the technical setup around it. That's fair, and it's precisely why ARC-AGI-3 deliberately tests pure model performance without external aids. Opus 5 hit its 30 percent under those same constraints and would likely score even higher inside Claude Code.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

