
NTU 曹子昂教授团队:破解 3D 标注成本难题,只需一张图片丨CVPR 2026
Quick Answer
NTU's PhysX-Anything model generates simulation-ready physical 3D assets from a single image, significantly improving the accuracy of geometric and physical properties.
Quick Take
It outperforms competitors like URDFormer and PhysXGen, achieving PSNR of 20.35 and reducing scale error from 43.44 to 0.30, enabling practical applications in robotics and AR/VR.
Key Points
- PhysX-Anything converts single images into interactive 3D assets for robotics training.
- Achieved PSNR of 20.35 and reduced scale error to 0.30, outperforming existing models.
- Supports applications in AR/VR, enabling realistic user interactions with virtual objects.
- Utilizes a novel token compression method, reducing representation size by 193 times.
- Developed by NTU's Ziang Cao, focusing on physical intelligence in 3D asset generation.
Source Excerpt
PhysX-Anythingt:可从一张照片自动生成可用于机器人训练的物理 3D资产。 作者丨郑佳美、樊天骄 编辑丨郑佳美 在生成式 AI 进入 3D 内容生产之后,行业最先解决的是“看起来像不像”的问题:一个模型能不能从文字或图片生成外观完整、纹理逼真、形状合理的 3D 物体。 但随着机器人、具身智能、数字孪生、AR / VR 和工业仿真的发展,真正制约应用落地的矛盾已经变了。 现实世界中的物体不是静态摆件,而是带有尺度、材料、重量、关节、摩擦、碰撞和功能关系的物理对象。 一个柜子不仅要有柜门,还要知道门轴在哪里、能向哪个方向打开;一副眼镜不仅要有镜框和镜腿,还要知道镜腿能绕哪个关节折叠;一个水龙头不仅要外形相似,还要能被旋转、能和机械手发生接触、能在仿真器里表现出合理运动。 换句话说,未来的 3D 生成如果只停留在“生成一个好看的模型”,就很难支撑机器人训练、交互式场景构建和真实物理仿真。 这正是当前 3D 资产生成面临的核心断层:视觉资产越来越容易生成,但仿真资产依然高度依赖人工建模和手动标注。 这个过程成本高、效率低,也很难规模化扩展到家庭、工厂、商场、医院等复杂真实场景。
因此,行业真正需要的不只是“图像到 3D”,而是“图像到可交互、可运动、可仿真的物理 3D 资产”。 在这种背景下,南洋理工大学曹子昂团队提出了《PhysX-Anything: Simulation-Ready Physical 3D Assets from Single Image》。 试图把单张真实图像直接转化为仿真可用的物理 3D 资产。 …
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from 雷峰网 AI
See more →
刚刚,GPT 5.6 发布会上,OpenAI 暴露了哪些 Agent 技术路线?
OpenAI's GPT 5.6 integrates ChatGPT and Codex, introducing a for complex task execution, with models Soul, Terra, and Luna for efficient workflow management. The release emphasizes task orchestration, contextual understanding, and robust security measures for enterprise applications.

