
Moonshot's Kimi K3 outperforms Fable 5 in frontend code but lags far behind in complex math
Quick Answer
Moonshot's Kimi K3 excels in frontend code with a score of 1,679, surpassing Fable 5's 1,631, marking a first for Chinese models.
Quick Take
However, it struggles in complex math, achieving only 39% accuracy on FrontierMath Tier 4, while competitors like OpenAI and Anthropic reach close to 90%.
Key Points
- Kimi K3 scores 1,679 in frontend benchmarks, leading the field.
- Fable 5 scores 1,631, trailing Kimi K3 in human preference ratings.
- Kimi K3 only achieves 39% accuracy on the hardest math tasks.
- Top Western models score around 90% on similar complex math benchmarks.
- This performance highlights disparities in AI capabilities across regions.
📖 Reader Mode
~1 min readMoonshot's AI model Kimi K3 is getting a lot of attention in the Western AI community. The big question is how close it actually gets to the best Western models. Two new data points paint a mixed picture. In the Code Arena: Frontend benchmark, which ranks models based on human preference ratings, Kimi K3 scores 1,679, beating Claude Fable 5 (1,631), GPT-5.6 Sol (1,618), and every other tested model by a wide margin. It's the first time a Chinese model has claimed the top spot on this benchmark.
The picture looks different for hard math. According to data from Epoch AI, Kimi K3 hits only about 39 percent accuracy on FrontierMath Tier 4, the benchmark's hardest expert-level math tasks. Models from OpenAI and Anthropic score close to 90 percent there in some cases.

— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

