
Anthropic's Opus 5 blows past Fable 5 and GPT-5.6 Sol on the benchmark designed to measure real intelligence
Quick Answer
Anthropic's Claude Opus 5 outperformed OpenAI's GPT-5.6 Sol with a score of 30.2% on the ARC-AGI-3 benchmark, showcasing superior reasoning capabilities.
Quick Take
It solved five previously unsolved environments, demonstrating significant advancements in autonomous exploration and planning.
Key Points
- Opus 5 leads -3 with a score of 30.2%, far surpassing GPT-5.6 Sol's 7.8%.
- It solved five previously unsolved environments, four at or above human level.
- Opus 5 scored 90.4% on ARC-AGI-2 and 97.5% on ARC-AGI-1.
- The model demonstrated novel behaviors, including translating tasks into algebraic notation.
- Independent tests on Witness showed narrower gains compared to ARC-AGI-3.
DeepSignal Analysis
What happened
Anthropic's Claude Opus 5 achieved a score of 30.2% on the ARC-AGI-3 benchmark, surpassing OpenAI's GPT-5.6 Sol, which scored 7.8%. Opus 5 solved five previously unsolved environments, demonstrating advancements in reasoning and autonomous exploration. It also performed well on older benchmarks, scoring 90.4% on ARC-AGI-2 and 97.5% on ARC-AGI-1.
Key evidence
- Claude Opus 5 scored 30.2% on the ARC-AGI-3 benchmark, significantly higher than OpenAI's GPT-5.6 Sol, which scored 7.8%.
- Opus 5 solved five previously unsolved environments, with four of them achieving human-level performance.
- On the ARC-AGI-2 benchmark, Opus 5 scored 90.4%, and on ARC-AGI-1, it reached 97.5%, matching previous top scores.
Why it matters
The performance of Claude Opus 5 on the ARC-AGI-3 benchmark indicates a notable advancement in AI reasoning capabilities, particularly in solving new tasks without prior exposure. This could have implications for the development of more autonomous AI systems that can adapt to unfamiliar environments. The ability to solve previously unsolved tasks suggests that Opus 5 may represent a significant step forward in the quest for artificial general intelligence (AGI).
Source Excerpt
Anthropic's Claude Opus 5 scored 30. 2 percent on -3, nearly quadrupling GPT-5. 6 Sol's previous record of 7. 8 percent. The benchmark's developers say the model independently formulated reflection equations, a behavior they had never seen from another model, and attribute to stronger logical reasoning.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

