Video Models as Native 4D Renderers: World-Grounded Conditioning from Animated Mesh
Quick Answer
The DAR model enhances video diffusion rendering by integrating camera motion and animated mesh conditions, achieving PSNR of 25.36 on the DAR-4D benchmark.
Quick Take
This approach outperforms existing methods by improving PSNR by up to 1.54 dB, demonstrating the effectiveness of tracking and world position in 4D rendering.
Key Points
- DAR model projects a neural 4D G-buffer from animated meshes.
- Achieves PSNR of 25.36 and SSIM of 0.917 on the DAR-4D benchmark.
- Improves PSNR by 1.54 dB over Wan2.2-Depth.
- Tracking and world position are crucial for effective 4D rendering.
- Depth alone reduces PSNR by 1.26–1.55 dB at every checkpoint.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Junhao Chen, Mingjin Chen, Henghaofan Zhang, Minglin Chen, Liaoyuan Fan, Boran Zhang, Saining Zhang, Mingze Sun, Hao Zhao, Ruqi Huang, Zhihao Li, Yufei Li
Abstract:Pretrained video diffusion models can act as renderers when the desired scene state is already specified by an animated mesh, a camera trajectory, and a reference image. This 4D generative rendering setting raises a representation question: what image-format condition lets a video backbone obey both camera motion and scene-internal animation? We propose DAR, a reference-guided renderer that extends Wan2.2 camera control from Plücker rays alone to a joint camera-plus-geometry interface. DAR projects a neural 4D G-buffer (tracking, world position, and normal) from the animated mesh and injects it through a widened control adapter while preserving the pretrained image-to-video prior. The central design choice is the pair of tracking and world position. Tracking identifies the persistent surface element that should carry appearance; world position gives its current scene-coordinate state; normal supplies local shape. Depth plus calibrated rays can recover 3D in principle, but depth is a camera-dependent chart in which camera and object motion are mixed. On the 68-case DAR-4D benchmark, LoRA DAR reaches PSNR 23.22, SSIM 0.895, and LPIPS 0.134, improving over off-the-shelf Wan2.2-Depth by 1.54 dB PSNR; a full fine-tune reaches PSNR 25.36 and SSIM 0.917. Matched ablations show that replacing world position by depth reduces PSNR by 1.26--1.55 dB at every checkpoint, supporting tracking+world-position correspondence as a practical 4D rendering condition.
| Comments: | 14 pages, 5 figures |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2608.00094 [cs.CV] |
| (or arXiv:2608.00094v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.00094 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Mingjin Chen [view email]
[v1]
Thu, 30 Jul 2026 16:55:28 UTC (7,530 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.