
刚刚,DeepSeek V4 系列更新,架构没变,Agent 能力为何大涨
Quick Answer
DeepSeek's V4-Flash model has been updated with enhanced agent capabilities through retraining, achieving benchmark scores of 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE.
Quick Take
While the architecture remains unchanged, the new Responses API facilitates better integration for developers.
Key Points
- V4-Flash-0731 achieved 82.7 on 2.1 and 54.4 on DeepSWE.
- No changes were made to the model architecture or parameter size.
- The update focuses on retraining and improved API integration for agents.
- DeepSeek has not disclosed specific retraining data or methods.
- Performance improvements are based on internal testing environments.
DeepSignal Analysis
What happened
DeepSeek has updated its V4-Flash model, enhancing agent capabilities through retraining while maintaining the same architecture. The model achieved scores of 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE. Additionally, the introduction of the Responses API aims to improve integration for developers.
Key evidence
- DeepSeek's V4-Flash model was updated without changing its architecture or parameter scale, focusing instead on retraining.
- The updated V4-Flash achieved a score of 82.7 on Terminal Bench 2.1 and 54.4 on DeepSWE, indicating improved performance in agent tasks.
- The new Responses API supports better integration for developers, facilitating communication between the model and various tools.
Why it matters
This update demonstrates that significant improvements in model performance can be achieved without altering the underlying architecture. The focus on retraining suggests a shift towards optimizing existing models rather than solely expanding their capabilities. However, the lack of transparency regarding the retraining data and methods raises questions about the reproducibility of these results.
What to watch
📖 Reader Mode
~2 min read
作者丨郑佳美
编辑丨马晓宁
刚刚,DeepSeek 更新了 DeepSeek-V4-Flash。
官方把它称为“正式版”,不过目前只是在 API 端上线公测。此次更新也只涉及 deepseek-v4-flash,V4-Pro API,以及 App 和网页端当前使用的模型,都没有随之更新。

但比起“正式版”这个名字,这次更新更值得关注的是另一句话:DeepSeek-V4-Flash-0731 与此前的 Preview 版本采用相同的模型结构和参数规模,只重新进行了后训练。


01
没有换架构,为什么能力还能提升?
大模型的训练其实可以被简单分成两个阶段。
第一个阶段是预训练。模型通过大量文本和代码学习知识,形成基础能力。这个阶段决定了模型大致有多聪明、知道多少东西。
第二个阶段是后训练。它更像是针对具体工作进行专项训练,让模型学会怎样回答问题、怎样调用工具,以及怎样按照要求完成任务。
这次 DeepSeek 没有重新设计模型架构,也没有扩大参数规模,而是在原有模型上重新进行后训练。模型的结构没有变化,但内部权重和行为方式会发生变化。
这有点像同一辆车没有更换发动机,却重新调整了变速箱、控制系统和驾驶策略。车本身没有变大,但在特定道路上的表现可能明显改善。
不过,DeepSeek 并没有公布这次后训练具体使用了哪些数据、奖励方法和训练流程。因此,目前只能确定它重新做了后训练,不能进一步断言性能提升究竟来自哪一种技术。

02
提升主要指向 Agent 能力
根据 DeepSeek 公布的结果,新版 V4-Flash 在 Terminal Bench 2.1 上获得 82.7,在 DeepSWE 上获得 54.4,在 DSBench-FullStack 上获得 68.7。

这些测试和传统的代码补全不同。传统代码测试通常只要求模型写出一段代码,而 Agent 测试要求模型真正进入一个工作环境。它需要读取文件、理解代码仓库、调用终端、修改程序、运行测试,再根据报错继续处理。
也就是说,模型不只是要知道答案,还要连续完成多个步骤。
不过,这些成绩也不能完全理解成模型自身的单独能力。DeepSeek 明确说明,部分代码 Agent 测试使用了尚未公开的 DeepSeek Harness,并采用了较高的推理档位;DSBench-FullStack 和 DSBench-Hard 还是 DeepSeek 自己的内部测试集。
因此,更准确的说法是:V4-Flash-0731 在 DeepSeek 公布的 Agent 运行环境中取得了明显提升,但它在第三方框架中的实际表现,还需要更多独立测试。

03
Responses API 解决的是接入问题
此次更新还有一个重要变化:V4-Flash 开始原生支持 Responses API,并针对 Codex 进行了适配。
Responses API 可以理解为一套更适合 Agent 的通信接口。它可以更统一地处理模型回复、推理状态、工具调用和执行结果。

它本身不会让模型突然变得更聪明,但可以减少模型接入 Codex 等 Agent 工具时的适配工作,让开发者更容易把模型放进真实的软件开发流程。雷峰网(公众号:雷峰网)
所以,这次更新其实包含两部分:一部分是重新后训练后的模型权重,另一部分是更适合 Agent 使用的接口和运行环境。
最终成绩由模型、推理预算、工具环境和 Agent 框架共同决定,很难只归功于其中某一项。

04
这次更新真正值得关注的是什么?
目前可以确定的是,DeepSeek 在没有改变 V4-Flash 模型结构和参数规模的情况下,通过重新后训练提高了官方 Agent 基准成绩,同时增强了对 Responses API 和 Codex 工作流的支持。
但现阶段还不能直接得出后训练已经可以替代模型扩容或者V4-Flash 已经全面领先其他 Agent 模型这样的结论。
一方面,V4-Flash-0731 的具体后训练变化尚未公开;另一方面,官方部分测试所使用的 Harness 和内部数据集还不具备完整的外部复现条件。
所以,这次更新不适合简单总结为DeepSeek 又发布了一个更强的模型。更准确的理解是,DeepSeek 正在验证一条更加务实的技术路线:当模型结构和参数规模保持不变时,能否通过重新后训练、工具接口适配和 Agent 运行框架,进一步提高模型完成真实任务的能力。雷峰网
DeepSeek 已经给出了官方成绩,但这条路线究竟能带来多大的实际提升,还要等待更多训练细节、公开工具和第三方测试。


上车,带你看遍全球 AI 顶会精华
可独家畅览:
专家演讲PPT
大会报告全文
热门论文解读
学术新星访谈

扫描上方二维码
或点击「阅读原文」关注专区。
雷峰网原创文章,未经授权禁止转载。详情见转载须知。
— Originally published at leiphone.com
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from 雷峰网 AI
See more →
刚刚,GPT 5.6 发布会上,OpenAI 暴露了哪些 Agent 技术路线?
OpenAI's GPT 5.6 integrates ChatGPT and Codex, introducing a for complex task execution, with models Soul, Terra, and Luna for efficient workflow management. The release emphasizes task orchestration, contextual understanding, and robust security measures for enterprise applications.

