LipSSD: Lipschitz-Constrained Single-Shot Detection for Adversarially Robust Object Detection
Quick Answer
LipSSD introduces a Lipschitz-constrained Single Shot MultiBox Detector, enhancing adversarial robustness in object detection.
Quick Take
It shows up to a 15-point improvement in mAP@50 on unseen attacks compared to traditional adversarially trained SSDs, while maintaining performance on safety-critical datasets like LARD and KITTI.
Key Points
- LipSSD is designed to be robust against adversarial attacks in object detection.
- Lipschitz constraints allow control over the accuracy-robustness trade-off via a single hyperparameter.
- Adversarially trained LipSSD outperforms classical SSD by up to 15 points on unseen attacks.
- The model maintains clean performance while improving robustness on datasets like LARD and KITTI.
- Architectural Lipschitz control offers a practical, attack-agnostic approach to enhance detector robustness.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Object detectors have many applications in safety-critical systems, but they are known to be sensitive to worst-case perturbations such as adversarial attacks, which limits their applicability in real-world scenarios. Compared with classification, adversarial robustness for object detection has received less attention, and existing methods are often tied to adversarial training, whose performance may not transfer across attacks, perturbation budgets, or architectures. In this work, we introduce Lipschitz-constrained variants of object detection architectures as robust-by-design alternatives to standard detectors. We validate this approach with LipSSD, a Lipschitz-constrained Single Shot MultiBox Detector (SSD), and provide a comprehensive study of its adversarial robustness using multiple white-box adversarial attacks and datasets. We first analyze the accuracyrobustness trade-off induced by Lipschitz constraints and show that it can be controlled through a single training hyperparameter. We then demonstrate that Lipschitzconstrained detectors are complementary to adversarial training: under the same training setup on the Pascal VOC dataset, adversarially trained LipSSD improves mAP@50 on unseen attacks by up to 15 points over classical adversarially trained SSD. Finally, we use more specific safety-critical datasets such as LARD and KITTI, and show that Lipschitz-constrained detectors can improve robustness while largely preserving clean performance. These results suggest that architectural Lipschitz control is a practical and attack-agnostic direction for improving the robustness of object detectors.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.06592 [cs.CV] |
| (or arXiv:2607.06592v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.06592 arXiv-issued DOI via DataCite |
Submission history
From: Corentin FRIEDRICH [view email] [via CCSD proxy]
[v1]
Mon, 6 Jul 2026 08:33:28 UTC (5,386 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.