
New Deepseek Flash model matches OpenAI's GPT-5.6 Luna at roughly 60 percent lower cost
Quick Answer
Deepseek's V4 Flash '0731' model rivals OpenAI's GPT-5.6 Luna, scoring 50 points and costing 60% less per task.
Quick Take
It shows significant improvements in agentic tasks and uses fewer tokens, while maintaining a similar architecture with 284 billion parameters.
Key Points
- Deepseek V4 Flash '0731' scores 50 points, just one point behind GPT-5.6 Luna.
- The new model costs about 60% less per task than OpenAI's offering.
- Improvements noted in agentic tasks and reduced hallucination rates.
- It uses 12% fewer tokens than its predecessor, enhancing efficiency.
- Model weights are available under MIT license on Hugging Face.
📖 Reader Mode
~2 min readDeepseek has released V4 Flash "0731," a major upgrade to its budget AI model. According to the Artificial Analysis Intelligence Index, the new version scores 50 points, ten more than the previous V4 Flash that launched in April 2026. That puts it just one point behind OpenAI's budget model GPT-5.6 Luna, but it costs about 60 percent less per task, even after OpenAI's 80 percent price cut. A big reason for the gap is Deepseek's 98 percent cache discount, well above the industry-standard 90 percent. The model also uses 12 percent fewer tokens than its predecessor.

The model improves across every tested category compared to the previous version, with the biggest gains in agentic tasks. On GDPval, a benchmark designed to test models on complex real-world office work, it climbs from 1,189 to 1,559 Elo points. It also hallucinates less often. The architecture stays the same: 284 billion total parameters, 13 billion active, with a one-million-token context window. The model weights are available under an MIT license on Hugging Face.
— Originally published at the-decoder.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from The Decoder
See more →
An AI model programmed nonstop for 19 days on a single MirrorCode task that cost $2,600 to run
Epoch AI's MirrorCode benchmark reveals Claude Opus 4.7 as the leader with a 56% solve rate, reconstructing a 16,000-line toolkit in 14 hours. Despite this, all models tested struggle with the most complex tasks, highlighting limitations in current AI capabilities. The single task consumed $2,600 over 19 days, raising questions about cost-effectiveness in AI development.

