RxBrain: Embodied Cognition Foundation Model with Joint Language-Visual Reasoning and Imagination
Quick Answer
RxBrain is an embodied cognition foundation model that integrates joint language-visual reasoning and imagination, enabling agents to connect high-level task reasoning with physical states.
Quick Take
It utilizes a unified Mixture-of-Transformers architecture for multimodal understanding and generation, showing promising results in continuous robot action generation without extensive pretraining.
Key Points
- RxBrain employs a unified Mixture-of-Transformers architecture for multimodal tasks.
- It integrates language and visual imagination for embodied planning.
- The model shows promising performance in real-robot action generation.
- RxBrain-Bench evaluates joint textual and visual components in planning.
- Automatic pipelines convert embodied videos into planning supervision.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Haotian Liang, Mingkang Chen, Yufei Huang, Yuchun Guo, Xiaomeng Zhu, Xiangli Shi, Kaixuan Wang, Yunxuan Mao, Weijie Zhou, Ling Chen, Shirong Zeng, Yueyu Long, Yuchen Si, Yajuan Zhu, Xingyu Zhou, Minghui Wang, Wanjia He, Xin Yang, Lingzhu Xiang, Zhiqing Liu, Bohan Ma, Xiran Huang, Tianshuo Yang, Zhiheng Liu, Xuantang Xiong, Zisheng Lu, Ping Luo, Yao Mu, Han Hu, Zhengyou Zhang
Abstract:Embodied cognition requires agents to connect high-level task reasoning with the physical states to be achieved. We introduce Hy-Embodied-RxBrain, an embodied cognition foundation model with joint language-visual reasoning and imagination. Unlike vision-language models that emphasize scene understanding and textual decision making, or generative world models that mainly predict future visual states, RxBrain represents embodied plans in a single planning sequence where language and visual imagination play complementary roles. Language provides the abstract structure of a plan, including task decomposition, planning primitives, constraints, temporal order, and decision logic, while visual imagination grounds this structure through world state prediction and joint subgoal planning, associating each planning step with intermediate and final physical states. RxBrain adopts a unified multimodal Mixture-of-Transformers architecture that supports language, image, and video understanding and generation within one model. To train this capability, we build an automatic pipeline that converts embodied videos into joint text-visual planning supervision by decomposing videos into planning steps and aligning them with visual state transitions. We further introduce RxBrain-Bench to evaluate whether models can represent embodied plans through joint textual and visual components rather than separate understanding or generation. Experiments show that RxBrain maintains embodied understanding and generation abilities, and produces plans with coupled textual reasoning, world state prediction, and joint subgoal planning. We also extend RxBrain to continuous robot action generation, where it shows promising real-robot performance without large-scale action-data pretraining. These results provide an initial step toward foundation models for embodied cognition.
| Subjects: | Artificial Intelligence (cs.AI); Robotics (cs.RO) |
| Cite as: | arXiv:2607.14187 [cs.AI] |
| (or arXiv:2607.14187v1 [cs.AI] for this version) | |
| https://doi.org/10.48550/arXiv.2607.14187 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Haotian Liang [view email]
[v1]
Wed, 15 Jul 2026 15:45:25 UTC (17,913 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.AI
See more →HOBA: Hierarchical On-Policy Bidding Agents for Adaptive Online Advertising
HOBA (Hierarchical On-policy Bidding Agents) is a novel hierarchical reinforcement learning framework that enhances online advertising bidding systems by improving adaptability and reducing hyperparameter tuning costs. It utilizes a for hyperparameter inference, a SARSA agent for expert model selection, and a dynamic expert pool for bid execution, achieving a +3.6% increase in target cost during large-scale deployment and outperforming state-of-the-art baselines on AuctionNet.