A Task-Driven Evaluation of UAV Detection and Tracking under Synthetic Fog
Quick Answer
This study introduces a task-driven evaluation framework for UAV detection and tracking under synthetic fog, revealing that fog significantly impairs performance due to increased missed detections.
Quick Take
It shows that training with fog-inclusive data enhances robustness, while restoration techniques are most effective when detectors are trained on clear imagery.
Key Points
- Synthetic fog is generated from clear-weather UAV images using monocular depth estimation.
- Fog inclusion during training improves detection robustness significantly.
- Restoration quality does not directly correlate with improved detection and tracking performance.
- Detection and tracking performance degrade substantially in foggy conditions.
- Multiple restoration methods were evaluated, including classical and CNN-based approaches.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Fog severely degrades the visibility of small unmanned aerial vehicles (UAVs) in skydominant, long-range imagery, reducing the reliability of downstream detection and tracking. This paper presents a task-driven evaluation framework that links depth-aware synthetic fog generation, image restoration, object detection, and tracking within a unified pipeline. Given the practical difficulty of collecting and annotating foggy UAV scenes, synthetic fog is generated from real clear-weather outdoor images containing UAV targets using monocular depth estimation and the atmospheric scattering model. Representative restoration methods from classical, convolutional neural network (CNN)-based, and transformer-based families are first compared, after which the selected restoration model is integrated into the downstream perception pipeline. Detection is evaluated under both clean-only and fog-inclusive training regimes using multiple detector variants, while tracking-by-detection is assessed on clean, foggy, and restored video sequences. Beyond image-level restoration metrics, the study evaluates how fog and restoration affect detection robustness and tracking performance. The results show that fog substantially degrades both detection and tracking, primarily through increased missed detections. Fog-inclusive training provides the most consistent improvement in robustness, whereas test-time restoration is most beneficial when the detector has been trained only on clean imagery. These findings show that restoration quality does not necessarily translate into proportional gains in downstream perception and therefore should be evaluated jointly with detection and tracking performance.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG); Image and Video Processing (eess.IV) |
| Cite as: | arXiv:2607.05467 [cs.CV] |
| (or arXiv:2607.05467v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.05467 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Vesal Ahsani [view email]
[v1]
Mon, 6 Jul 2026 07:06:37 UTC (4,824 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.