muSync-GS: Physics-Synchronized Driving Video Synthesis for Weather and Geometric Road Hazards
Quick Answer
This paper shows that muSync-GS introduces a physics-synchronized framework for synthesizing driving videos under adverse weather and road conditions, achieving RMSEs of 0.0273 m/s for speed and 0.0590 degrees for pitch.
Quick Take
This model enhances the realism of autonomous driving simulations by coupling road conditions with vehicle dynamics, addressing challenges in data collection for rare scenarios.
Key Points
- Achieves mean RMSEs of 0.0273 m/s for speed and 0.0590 degrees for pitch.
- Integrates precipitation-derived road conditions with tire friction and vehicle dynamics.
- Synchronizes ego-camera motion with controllable scene edits.
- Demonstrates effectiveness across 12 CarSim cases with varying parameters.
- Addresses safety-critical scenarios in autonomous driving data collection.
DeepSignal Analysis
What happened
The muSync-GS framework was developed to synthesize driving videos that accurately reflect adverse weather and road conditions. It integrates vehicle dynamics with road conditions to enhance the realism of simulations, achieving low RMSE values for speed and pitch.
Key evidence
- muSync-GS achieves a mean case-wise RMSE of 0.0273 m/s for speed and 0.0590 degrees for pitch, indicating high accuracy in vehicle response simulation.
- The framework couples road appearance and tire friction through precipitation-derived road-surface conditions, addressing limitations in existing video-generation methods.
- The model was tested on 12 held-out CarSim cases, demonstrating its ability to synchronize vehicle dynamics with controllable scene edits.
Why it matters
This framework addresses a critical gap in autonomous driving simulations by providing a method to generate realistic driving scenarios under challenging conditions. The ability to simulate rare and safety-critical situations can enhance the training of autonomous systems, potentially improving their safety and reliability in real-world applications.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAbstract:High-quality driving data are essential for autonomous-driving systems and generative world models. However, rare and safety-critical scenarios involving adverse weather, braking under low tire--road friction, and uneven road geometry are costly and risky to collect at scale. Existing video-generation and 3D Gaussian editing methods can modify weather appearance or road geometry, but typically do not couple these edits with tire--road interaction and vehicle dynamics. As a result, an edited video may retain its original trajectory even when the modified road condition should alter braking, wheel slip, load transfer, and ego-camera motion. We present muSync-GS, a physics-synchronized framework for driving video synthesis under adverse-weather and road-elevation hazards. A precipitation-derived road-surface condition jointly controls road appearance and tire friction, while a shared road-elevation profile drives both visible road-geometry editing and axle excitation. A calibrated vehicle model predicts speed, slip ratio, normal loads, and pitch for constructing the ego-camera trajectory and synchronized physical annotations. On 12 held-out CarSim cases spanning precipitation levels, brake inputs, and road-profile parameters, the model achieves mean case-wise RMSEs of 0.0273 m/s for speed, 0.0590 degrees for pitch, 0.0101 for slip ratio, and 26.61 N for per-wheel normal load. Together with the reconstructed-scene experiments, these results show that muSync-GS accurately reproduces vehicle responses under held-out controls while synchronizing them with controllable scene edits and ego-camera motion.
| Comments: | 42 pages, 14 figures; includes an appendix |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2608.04412 [cs.CV] |
| (or arXiv:2608.04412v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2608.04412 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Yang Chen [view email]
[v1]
Wed, 5 Aug 2026 03:47:18 UTC (9,835 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.