RealVDeblur: One-Step Diffusion for Generalizable Real-World Video Deblurring
Quick Answer
RealVDeblur introduces an efficient generative framework for video deblurring, leveraging a large-scale blur synthesis pipeline and a one-step diffusion generator.
Quick Take
It achieves strong perceptual quality and temporal consistency in unseen videos, enhancing robustness for applications like mobile imaging and 3D reconstruction under severe motion blur.
Key Points
- Constructs a large-scale blur synthesis pipeline using 3D Gaussian Splatting assets.
- Utilizes a video diffusion prior for effective restoration under diverse conditions.
- Employs a one-step generator for efficient long video processing.
- Demonstrates improved robustness in downstream 3D reconstruction tasks.
- Achieves strong perceptual quality and temporal consistency on real-world benchmarks.
Paper Resources
📖 Reader Mode
~2 min readAuthors:Renbiao Jin, Mingxin Yang, Yutian Chen, Junhao Zhuang, Xin Cai, Mulin Yu, Linning Xu, Wenxian Yu, Danping Zou, Shi Guo, Tianfan Xue
Abstract:Real-world video deblurring remains challenging due to diverse motion patterns, complex degradations, and the scarcity of realistic training data, yet robust restoration is critical for downstream pipelines such as mobile imaging and 3D reconstruction. This work presents \textbf{RealVDeblur}, an efficient generative framework designed to improve in-the-wild robustness under diverse real capture conditions. First, a large-scale, physically grounded blur synthesis pipeline is constructed from scene-level 3D Gaussian Splatting (3DGS) assets and high-frame-rate videos, providing realistic training data covering both camera-induced and object-motion blur. Second, a video diffusion prior is leveraged for restoration; to better accommodate frame-dependent blur variations, temporal compression in the VAE is disabled and a frame-wise encoding scheme is adopted. For practical deployment on long videos, multi-step diffusion sampling is distilled into an efficient one-step generator, and a training-free Temporal Window Mask stabilizes inference beyond the training horizon with constant memory usage. Extensive experiments on diverse real-world benchmarks demonstrate strong perceptual quality, semantic fidelity, and temporal consistency on unseen videos, as well as improved robustness in downstream 3D reconstruction under severe motion blur. Project page: this https URL
| Comments: | Project page with code: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.20628 [cs.CV] |
| (or arXiv:2607.20628v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.20628 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Renbiao Jin [view email]
[v1]
Wed, 22 Jul 2026 18:01:03 UTC (5,510 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.