
Claude Fable 5 outpaces GPT-5.5 by 13 points on FrontierMath's toughest problems
Quick Answer
Anthropic's Claude Fable 5 achieves 88% accuracy on FrontierMath's toughest problems, surpassing OpenAI's GPT-5.5 by 13 points at 75%.
Quick Take
This significant improvement marks a leap from Opus 4.5's sub-10% performance in early 2026, highlighting the rapid advancement in AI math capabilities.
Key Points
- Claude Fable 5 scores 88% on FrontierMath's hardest problems.
- OpenAI's GPT-5.5 achieves 75% accuracy on the same benchmark.
- Fable 5 outperforms GPT-5.5 by a notable 13 points.
- Opus 4.5 had less than 10% accuracy in early 2026.
- AI math performance is improving at an accelerating pace.
📖 Reader Mode
~1 min readAnthropic's new model, Claude Fable 5, posts top scores on the FrontierMath benchmark. According to Epoch AI, Fable 5 hits 87 percent accuracy on tiers 1 through 3 and 88 percent on the hardest tier 4 (v2).

Anthropic's models are getting dramatically better at math in a short span of time. As recently as early 2026, predecessor model Opus 4.5 scored below 10 percent on tier 4. OpenAI's GPT-5.5 reaches about 75 percent on the same tier, well behind Fable 5, although GPT-5.6 is already in the making.
All models were tested on Epoch AI's standard scaffold with maximum reasoning effort. FrontierMath is widely considered one of the toughest benchmarks for AI math reasoning. These math gains aren't just in benchmarks, real-world examples keep stacking up. Most recently, an OpenAI model solved a longstanding Erdős problem; so did Claude Mythos.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

