FogDrive: A Multi-Modal Synthetic Driving Dataset for Perception under Graded Fog
Quick Answer
FogDrive is a new multi-modal synthetic driving dataset designed for evaluating perception in foggy conditions, featuring 660 scenes and 133k annotated frames from synchronized cameras and LiDAR.
Quick Take
It establishes benchmarks using architectures like TransFusion and YOLOv8-m, revealing that mixed fog training improves 3D bounding box accuracy without extra data costs.
Key Points
- FogDrive includes 660 scenes with 133k fully annotated frames across multiple sensors.
- Fog modeled using Koschmieder and Beer-Lambert laws at three visibility densities.
- Achieved 95.1% precision and over 99% recall for vehicle detection within 40m.
- Benchmarks established with TransFusion, BEVFusion, and YOLOv8-m for 3D fusion and 2D restoration.
- Dataset will be open-sourced to promote multi-modal research.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Perception under adverse weather remains a critical bottleneck for reliable autonomous driving, yet existing benchmarks lack the systematic multi-modal alignments needed to evaluate robust sensor fusion. Real-world weather datasets suffer from uncontrolled collection and single-level, uncalibrated conditions, while synthetic alternatives either target camera-only restoration or lack the paired clean-and-foggy structure needed to benchmark "defog-then-detect" pipelines. We present FogDrive, a rigorously calibrated, multi-modal autonomous-driving dataset bridging data-centric engineering and robust machine learning. Built with the CARLA simulator, FogDrive contains 660 scenes (~133k fully annotated frames, 50:50 day/night) across four synchronized cameras (RGB, depth, semantic segmentation), a LiDAR and semantic-LiDAR pair, and front radar. Physically consistent fog is modeled independently on camera channels (Koschmieder model) and LiDAR channels (Beer-Lambert law) at three calibrated visibility densities (160m, 100m, 50m). Every scene ships in four matched variants (clean plus three graded fog levels) with cross-calibrated 2D and 3D bounding boxes. A semantic-segmentation-based quality audit over 8k images validates annotations at 95.1% precision and over 99% recall for vehicles within 40m. We establish baseline benchmarks with state-of-the-art architectures (TransFusion, BEVFusion, YOLOv8-m) across two paradigms: 3D multi-modal fusion and 2D image restoration. These yield critical data-centric insights: mixing multi-density fog during training tightens 3D bounding-box geometry without added data-scaling cost, while in 2D pipelines image-quality metrics (PSNR, SSIM) prove poor predictors of downstream detection performance. FogDrive will be fully open-sourced alongside our data-generation framework to accelerate robust, multi-modal research.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.22698 [cs.CV] |
| (or arXiv:2607.22698v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.22698 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Vansh Panwar [view email]
[v1]
Sat, 18 Jul 2026 06:53:56 UTC (2,134 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.