
Alibaba’s open-weight Qwen3.8-Max takes on long-horizon AI tasks with 2.4 trillion parameters
Quick Answer
Alibaba's Qwen3.8-Max, with 2.4 trillion parameters, excels in long-horizon AI tasks, outperforming human teams in coding and chip design.
Quick Take
The model autonomously completed complex tasks, achieving a balance of 416,252 yuan in a simulated e-commerce scenario, significantly surpassing its predecessor Qwen3.7-Max.
Key Points
- Qwen3.8-Max features 2.4 trillion parameters, focusing on long-duration tasks.
- The model autonomously built a command-line tool over 16 days with 265 commits.
- In chip design, it reduced gate count from 8,298 to 678, optimizing efficiency.
- Qwen3.8-Max achieved a balance of 416,252 yuan in an e-commerce simulation.
- It processes multimodal tasks, handling documents over 200 pages and videos over 100 hours.
DeepSignal Analysis
What happened
Alibaba has introduced its Qwen3.8-Max model, featuring 2.4 trillion parameters, designed for long-duration AI tasks. The model autonomously completed complex coding and chip design tasks, significantly outperforming its predecessor and human teams in various benchmarks.
Key evidence
- Qwen3.8-Max autonomously built a command-line tool over 16 days, completing 265 commits and 151 issues without human intervention.
- In a simulated e-commerce scenario, Qwen3.8-Max achieved a balance of 416,252 yuan, more than 2.5 times its predecessor Qwen3.7-Max's performance.
- The model's performance on benchmarks like PaperBench reached a score of 93, the highest in comparison, although these results are self-reported and await independent verification.
Why it matters
The release of Qwen3.8-Max highlights Alibaba's advancements in AI, particularly in handling complex tasks over extended periods. This model's capabilities could influence the competitive landscape in AI development, especially against other models like Moonshot AI's Kimi K3. The focus on autonomous task completion may set new standards for AI applications in various industries.
What to watch
Source Excerpt
Alibaba's new flagship model Qwen3. 8-Max is built to handle complex tasks on its own over days at a time, from reproducing research papers to designing chips autonomously. The team plans to release the weights next week.
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

