
Qwen3.8 Max catches Claude Opus 4.8 but Kimi K3 still scores higher for 25 percent less
Quick Answer
Alibaba's Qwen3.8 Max scores 56 on the AI Index, matching Claude Opus 4.8 but lagging behind Kimi K3 at 57.
Quick Take
Despite lower token costs, its performance is hampered by increased task steps and a higher hallucination rate, leading to a cost of $1.14 per task, more than double that of Qwen3.7 Max.
Key Points
- Qwen3.8 Max scores 56, up from 46 in Qwen3.7 Max.
- Kimi K3 outperforms Qwen3.8 Max while being 25% cheaper.
- Qwen3.8 Max requires 64 steps per task, significantly more than previous models.
- Hallucination rate increased from 23% to 40% in Qwen3.8 Max.
- Cost per task for Qwen3.8 Max is $1.14, more than double Qwen3.7 Max.
DeepSignal Analysis
What happened
Alibaba's Qwen3.8 Max achieved a score of 56 on the AI Index, matching Claude Opus 4.8 but falling short of Kimi K3's score of 57. Although Qwen3.8 Max has lower token costs, its performance is hindered by a higher number of task steps and an increased hallucination rate, resulting in a cost of $1.14 per task, which is more than double that of its predecessor, Qwen3.7 Max.
Key evidence
- Qwen3.8 Max scored 56 on the AI Index, a significant increase from Qwen3.7 Max's score of 46.
- The cost per task for Qwen3.8 Max is $1.14, which is more than double the $0.53 cost of Qwen3.7 Max.
- The hallucination rate for Qwen3.8 Max increased from 23% to 40%, indicating a higher tendency to guess rather than admit lack of knowledge.
Why it matters
The performance of Qwen3.8 Max, despite its improved score, raises concerns about its efficiency due to the increased number of steps and hallucination rate. The higher cost per task could impact its competitiveness against models like Kimi K3, which offers a better price-to-performance ratio. This situation highlights the trade-offs between thoroughness and efficiency in AI model development.
What to watch
📖 Reader Mode
~2 min readAlibaba's Qwen3.8 Max scores 56 on the Artificial Analysis Intelligence Index, a 10-point jump over Qwen3.7 Max (46). According to Artificial Analysis, that puts it on par with Claude Opus 4.8 and ahead of GLM-5.2 (51), but behind Kimi K3 (57), which also runs 25 percent cheaper.
On GDPval-AA, a benchmark for work-related tasks, Qwen jumps 468 Elo points to 1,739, passing Kimi K3 (1,685). Only Claude Opus 5 (1,852) scores higher. The catch is how it gets there. Qwen3.8 Max needs 64 steps per task instead of 14, and input tokens grew 15x because the test resends the full conversation history to the model at each step.

The model works more thoroughly but runs slower and costs more. Alibaba's price-to-performance ratio takes a hit despite lower token prices (input dropped from $2.50 to $2.00 per million tokens, output from $7.50 to $6.00, and cache hits from $0.50 to $0.25). A single task in the Intelligence Index now costs $1.14, more than double Qwen3.7 Max ($0.53). Kimi K3 scores one point higher at just $0.86 per task, and GLM-5.2 comes in at $0.57.
There are also regressions compared to the previous version. AA-LCR dropped 2 points, a test that checks whether a model can correctly pull together information from very long texts. AA-Omniscience fell 10 points, measuring whether a model answers knowledge questions correctly or honestly admits it doesn't know. The accuracy rate stays around 31 percent, but the hallucination rate jumped from 23 to 40 percent. Qwen3.8 Max guesses far more often instead of saying it doesn't know.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

