
OpenAI claims GPT-5.6 Sol beats Opus 5 on ARC-AGI-3 with its latest API and two additional settings
Quick Answer
OpenAI's GPT-5.6 Sol outperforms Anthropic's Claude Opus 5 on the ARC-AGI-3 benchmark, scoring 38.3% versus Opus 5's 30.2%.
Quick Take
However, OpenAI utilized its own API settings, raising concerns about fairness in testing conditions. co-founder François Chollet acknowledges the complexity of benchmark setups and the need for clear reporting on settings used.
Key Points
- GPT-5.6 Sol scores 38.3% on ARC-AGI-3, surpassing Opus 5's 30.2%.
- OpenAI used its own Responses API with 'Retained Reasoning' and 'Compaction' settings.
- Official testing harness resulted in a low score of 7.8% for GPT-5.6 Sol.
- Chollet emphasizes the importance of clear reporting on testing setups.
- Discussions between ARC Prize and OpenAI continue regarding model testing standards.
DeepSignal Analysis
What happened
OpenAI's GPT-5.6 Sol achieved a score of 38.3% on the ARC-AGI-3 benchmark, surpassing Anthropic's Claude Opus 5, which scored 30.2%. However, OpenAI employed its own API settings, raising questions about the fairness of the comparison. François Chollet, co-founder of the ARC Prize, highlighted the importance of using standardized test setups to ensure equitable evaluations.
Key evidence
- OpenAI's GPT-5.6 Sol scored 38.3% on ARC-AGI-3, while Anthropic's Claude Opus 5 scored 30.2%.
- In the official test harness, GPT-5.6 Sol only scored 7.8% due to the loss of reasoning after each action.
- François Chollet noted that different settings among providers could create parity issues, but transparency in reporting is essential.
Why it matters
The results from OpenAI and Anthropic highlight the ongoing competition in AI model performance. However, the use of different testing setups complicates direct comparisons. Chollet's comments suggest that while OpenAI's approach may yield higher scores, it raises concerns about the validity of such benchmarks. Clear reporting of settings used is crucial for the integrity of AI evaluations.
What to watch
📖 Reader Mode
~2 min readUpdate:
ARC Prize co-founder François Chollet responded to OpenAI's results by distinguishing between two kinds of test setups. Harnesses "custom-made to solve the benchmark or that contain knowledge about the benchmark format" are off limits, he said. General-purpose API settings "that were not developed for ARC-AGI-3 and that are available to all API users" are fair game. In effect, Chollet is conceding that ARC Prize's own GPT-5.6 Sol score put OpenAI at a disadvantage.
He noted that ARC Prize has had "a lot of back and forth with OpenAI about how to best test their models, especially with regard to compaction," and welcomed the company "starting to figure out the answer." Different providers using different settings does create "a potential parity issue," Chollet said, but he considers that acceptable "as long as the settings and the cost are clearly reported."
Original article:
OpenAI says it can keep up on ARC-AGI-3. After Anthropic's Claude Opus 5 quadrupled the record score on the logic benchmark, OpenAI is now showing that GPT-5.6 Sol hits 38.3 percent with two API settings, beating Opus 5's 30.2 percent.

OpenAI isn't using the official test environment, though. Instead, it runs GPT-5.6 Sol through its own Responses API with "Retained Reasoning," which keeps the model's chain of thought between steps, and "Compaction," which summarizes old context instead of truncating it. In the official harness, GPT-5.6 Sol scored just 7.8 percent because the model's reasoning gets discarded after each action.
OpenAI argues that benchmarks never measure just the model but also the technical setup around it. That's true, and ARC-AGI-3 is designed to test pure model performance. The official ARC scores use a standardized approach without provider-specific settings to ensure fair comparisons, ARC Prize said in response to OpenAI's results. The sticking point is whether ARC Prize used an older "OpenAI-style completions API" that lacked features the Claude API already offered, which would make the comparison unfair to OpenAI.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

