
OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings
Quick Answer
OpenAI's GPT-5.6 Sol outperforms Anthropic's Claude Opus 5 on the ARC-AGI-3 benchmark, scoring 38.3% versus Opus 5's 30.2%.
Quick Take
However, OpenAI utilized its own API settings, raising concerns about fairness in testing conditions. co-founder François Chollet acknowledges the complexity of benchmark setups and the need for clear reporting on settings used.
Key Points
- GPT-5.6 Sol scores 38.3% on ARC-AGI-3, surpassing Opus 5's 30.2%.
- OpenAI used its own Responses API with 'Retained Reasoning' and 'Compaction' settings.
- Official testing harness resulted in a low score of 7.8% for GPT-5.6 Sol.
- Chollet emphasizes the importance of clear reporting on testing setups.
- Discussions between ARC Prize and OpenAI continue regarding model testing standards.
DeepSignal Analysis
What happened
OpenAI's GPT-5.6 Sol achieved a score of 38.3% on the ARC-AGI-3 benchmark, surpassing Anthropic's Claude Opus 5, which scored 30.2%. However, OpenAI employed its own API settings, raising questions about the fairness of the comparison. François Chollet, co-founder of the ARC Prize, highlighted the importance of using standardized test setups to ensure equitable evaluations.
Key evidence
- OpenAI's GPT-5.6 Sol scored 38.3% on ARC-AGI-3, while Anthropic's Claude Opus 5 scored 30.2%.
- In the official test harness, GPT-5.6 Sol only scored 7.8% due to the loss of reasoning after each action.
- François Chollet noted that different settings among providers could create parity issues, but transparency in reporting is essential.
Why it matters
The results from OpenAI and Anthropic highlight the ongoing competition in AI model performance. However, the use of different testing setups complicates direct comparisons. Chollet's comments suggest that while OpenAI's approach may yield higher scores, it raises concerns about the validity of such benchmarks. Clear reporting of settings used is crucial for the integrity of AI evaluations.
What to watch
Source Excerpt
OpenAI counters Anthropic's -3 record: GPT-5. 6 Sol scores 38. 3 percent, but only with its own API features instead of the official test setup, where the model landed at 7. 8 percent. ARC Prize claims its test environment is provider-neutral, but may have used an outdated API that skewed the comparison with Opus 5.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

