
让机器人行动更有依据:复旦等提出 GuidedVLA,提升 VLA 可控可解释能力
Quick Answer
GuidedVLA enhances the Visual-Language-Action (VLA) model by explicitly guiding action generation through task-relevant factors, achieving a success rate of 75.4% on the LIBERO-Plus benchmark, up from 68.2%.
Quick Take
This model improves interpretability and control, crucial for complex robotic tasks in dynamic environments.
Key Points
- GuidedVLA introduces Object, Skill, and Depth Heads for better task focus.
- Achieved a 90.63% success rate on RoboTwin 2.0, up from 77.38%.
- Automated labeling reduced annotation time from 43.5 minutes to 4 minutes.
- Model's interpretability allows for easier diagnosis of task failures.
- Research collaboration includes Fudan University and Shanghai Jiao Tong University.
Source Excerpt
GuidedVLA:以目标、阶段和空间约束,重塑 VLA 动作生成过程。 作者丨郑佳美 编辑丨马晓宁 机器人要进入更复杂的真实环境,真正的难点已经超出“能不能完成一个动作”。 更关键的问题是:当桌面变得杂乱、光照发生变化、任务步骤变长,或者目标物体变得透明、难以定位时,机器人能否稳定判断自己该看哪里、该做哪一步、空间位置是否准确。 这也是视觉-语言-动作模型(VLA)正在面对的核心挑战。 VLA 可以让机器人根据图像观测和语言指令生成动作,但在很多端到端训练框架中,动作生成过程仍然高度隐式。 模型给出了动作,却很难解释它依赖了哪些线索。 对真实机器人来说,可控可解释已经成为走向复杂任务的重要基础。 只有知道机器人为什么这样行动,研究者和工程团队才更容易诊断失败、改进模型,并把系统带到更多变化场景中。 围绕这一问题,复旦大学可信具身智能研究院联合上海交通大学、香港大学 OpenDriveLab 等机构提出了 GuidedVLA。 该工作已被 Robotics: Science and Systems(RSS)2026 接收,并开放了论文、项目主页、代码、模型权重和数据集。
GuidedVLA 的核心思路可以概括为一句话:在 VLA 的动作生成中加入显式引导,把任务相关因素拆成更清晰、更可检查的分工。 01机器人为什么需要“有依据地行动”过去几年,VLA 成为具身智能研究中的重要路线。 它把视觉理解、语言指令和动作生成连接起来:机器人看到环境,理解任务,再输出下一步动作。 这条路线的优势很明显。 模型结构更统一,训练方式更简洁,也更容易吸收视觉语言模型中的知识。 …
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from 雷峰网 AI
See more →
刚刚,GPT 5.6 发布会上,OpenAI 暴露了哪些 Agent 技术路线?
OpenAI's GPT 5.6 integrates ChatGPT and Codex, introducing a for complex task execution, with models Soul, Terra, and Luna for efficient workflow management. The release emphasizes task orchestration, contextual understanding, and robust security measures for enterprise applications.

