Fluid-Gen-Zero: Grounding Pretrained Video Generators in Physics without Training
Quick Answer
Fluid-Gen-Zero introduces a training-free framework for physics-aware video generation, improving simulation alignment by 26.7%-81.5% and fluid flow endpoint error by 67.9%-84.0%.
Quick Take
It leverages pretrained video generators and a physics simulator, outperforming existing methods in human preference studies, with 90.4%-94.2% favoring it over simulation-based approaches.
Key Points
- Fluid-Gen-Zero decouples physical reasoning from appearance synthesis.
- It employs a two-level agentic workflow for video generation.
- The framework shows significant improvements in object trajectory and fluid errors.
- Human preference studies indicate a strong favor for Fluid-Gen-Zero over alternatives.
- Code and data will be released upon acceptance of the paper.
Paper Resources
📖 Reader Mode
~2 min readAbstract:We present Fluid-Gen-Zero, a training-free framework for physics-aware fluid-object interaction video generation that decouples physical reasoning from appearance synthesis. Our key insight is to delegate motion dynamics to a physics simulator while preserving the appearance modeling capacity of pretrained video generators. We bridge these two domains through a two-level agentic workflow: generation-time planning, where a vision-language model (VLM) agent interprets intent and the simulation rollout to organize generation clips, and latent-space guidance, which injects simulation signals into denoising through region-aware latent wrapping. This plug-and-play design is compatible with current video foundation models. We further introduce a benchmark for fluid-object interaction video generation. Across Tora (CogVideoX-based), VACE and WanMove (Wan-based), Fluid-Gen-Zero consistently improves simulation alignment, reducing object trajectory error by 26.7%-81.5% and fluid fEPE (fluid flow endpoint error) by 67.9%-84.0%, while largely preserving perceptual quality. In a human preference study, raters favor Fluid-Gen-Zero in 55.1%-74.4% of same-backbone comparisons across three backbones, and in 90.4%-94.2% of comparisons against simulation-based methods. Code and data will be released upon acceptance.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Graphics (cs.GR) |
| Cite as: | arXiv:2610.10984 [cs.CV] |
| (or arXiv:2610.10984v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2610.10984 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Hong Huang [view email]
[v1]
Wed, 7 Oct 2026 23:16:32 UTC (10,404 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.