REVIVE: A Multi-Modal Framework for Vandalism Detection and Recovery in Autonomous Vehicles
Quick Answer
The REVIVE framework enhances vandalism recovery in autonomous vehicles by integrating binary detection, multi-class pattern identification, and EfficientNet-based segmentation, achieving a recall restoration from 0.588 to 0.967 using direct pixel replacement.
Quick Take
Stable Diffusion offers variable reconstruction performance, while a quality gate ensures downstream detection maintains or improves upon unrecovered baselines.
Key Points
- REVIVE integrates binary detection and multi-class pattern identification for vandalism recovery.
- Direct pixel replacement restores object-detection recall to 0.967 under aligned-reference conditions.
- Stable Diffusion shows variable reconstruction performance with SSIM ranging from 0.667 to 0.867.
- Quality gate filters recovered candidates, improving recall from 0.304 to 0.608.
- LaMa, Telea, and Navier-Stokes improve similarity but limit downstream detection recovery.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Autonomous vehicles (AVs) face increasing threats from vandalism-induced occlusion attacks (VOAs) that compromise camera-based perception. While detection frameworks can identify vandalized images, restoring camera-stream utility after physical occlusion remains underexplored. This paper presents present the Recovery and Enhancement of Vandalized Images for Vision Excellence (REVIVE) framework, a vandalism recovery pipeline integrating: (1) binary VOA detection, (2) multi-class VOA pattern identification, (3) EfficientNet-based U-Net segmentation, and (4) type-aware recovery using Bootstrapping Language-Image Pre-training (BLIP)-guided Stable Diffusion inpainting, direct pixel replacement, or adaptive median filtering. Stable Diffusion shows variable reconstruction performance (per-pattern SSIM 0.667-0.867, PSNR 15.4-26.7dB) across VOA patterns, while aligned direct pixel replacement achieves near-identical reconstruction under the aligned-reference condition. On 500 tracked clean/vandalized image pairs, unrecovered VOAs reduce YOLOv8l object-detection recall to 0.588, while direct pixel replacement restores recall to 0.967 and F1-score to 0.970 under that aligned-reference condition. LaMa, Telea, and Navier-Stokes baselines improve image similarity but provide more limited downstream detection recovery, and Stable Diffusion is treated as an asynchronous recovery branch subject to a quality gate rather than a blocking real-time perception step. We evaluate a reference-available quality gate that filters recovered candidates before downstream use: without it, type-aware routing degrades per-image recall to 0.304, whereas with it, recall returns to 0.608, at or above the unrecovered baseline, ensuring the forwarded stream is never worse than the unrecovered frame. REVIVE therefore, provides a structured recovery framework from VOAs in AVs.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Machine Learning (cs.LG) |
| Cite as: | arXiv:2607.05649 [cs.CV] |
| (or arXiv:2607.05649v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.05649 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Tapadhir Das [view email]
[v1]
Mon, 6 Jul 2026 21:25:41 UTC (2,579 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.