
刚刚,DeepSeek V4 系列更新,架构没变,Agent 能力为何大涨
Quick Answer
DeepSeek's V4-Flash model has been updated with enhanced agent capabilities through retraining, achieving benchmark scores of 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE.
Quick Take
While the architecture remains unchanged, the new Responses API facilitates better integration for developers.
Key Points
- V4-Flash-0731 achieved 82.7 on 2.1 and 54.4 on DeepSWE.
- No changes were made to the model architecture or parameter size.
- The update focuses on retraining and improved API integration for agents.
- DeepSeek has not disclosed specific retraining data or methods.
- Performance improvements are based on internal testing environments.
DeepSignal Analysis
What happened
DeepSeek has updated its V4-Flash model, enhancing agent capabilities through retraining while maintaining the same architecture. The model achieved scores of 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE. Additionally, the introduction of the Responses API aims to improve integration for developers.
Key evidence
- DeepSeek's V4-Flash model was updated without changing its architecture or parameter scale, focusing instead on retraining.
- The updated V4-Flash achieved a score of 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE, indicating improved performance in agent tasks.
- The new Responses API supports better integration for developers, facilitating communication between the model and various tools.
Why it matters
This update demonstrates that significant improvements in model performance can be achieved without altering the underlying architecture. The focus on retraining suggests a shift towards optimizing existing models rather than solely expanding their capabilities. However, the lack of transparency regarding the retraining data and methods raises questions about the reproducibility of these results.
What to watch
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from 雷峰网 AI
See more →
刚刚,GPT 5.6 发布会上,OpenAI 暴露了哪些 Agent 技术路线?
OpenAI's GPT 5.6 integrates ChatGPT and Codex, introducing a for complex task execution, with models Soul, Terra, and Luna for efficient workflow management. The release emphasizes task orchestration, contextual understanding, and robust security measures for enterprise applications.

