DiffAttack: Evasion Attacks Against Face Recognition via Latent Diffusion Models
Quick Answer
This paper shows that The DiffAttack framework utilizes latent diffusion models for adversarial face generation, achieving an 84.86% average attack success rate across multiple face recognition systems, significantly outperforming traditional methods.
Quick Take
This approach enhances transferability, surpassing noise-based methods by 15.28% and semantic-based techniques by 5.21% on datasets like FFHQ and CelebA-HQ.
Key Points
- DiffAttack leverages latent diffusion models for enhanced adversarial face generation.
- Achieves an average attack success rate of 84.86% on face recognition models.
- Outperforms traditional noise-based methods by over 15.28% in transferability.
- Demonstrates superior performance on FFHQ and CelebA-HQ datasets.
- Addresses limitations of existing adversarial methods in facial biometrics.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Facial biometric identification relies on the distinctiveness of user attributes within a high-dimensional embedding space. However, the decision boundaries of deep face recognition (FR) systems are often sufficiently narrow that they can be conflated, rendering the models vulnerable to adversarial attacks. In such scenarios, the FR system fails to distinguish between an authentic source and a meticulously crafted adversarial face. Existing adversarial methods targeting facial biometrics are limited in both performance and their ability to generate high-quality images that are imperceptible to humans. Moreover, these methods often fail when the source and target images belong to different demographic groups or genders. To address these limitations, we present a novel approach for adversarial face generation via latent-space optimization. We leverage latent diffusion models directly to guide generation toward target identity embeddings, as measured by a face recognition model. Our proposed \textbf{DiffAttack} framework has been evaluated on standard benchmarks, such as the FFHQ and CelebA-HQ datasets. DiffAttack significantly outperforms existing adversarial techniques, achieving a high average attack success rate of 84.86% across multiple face recognition models (e.g., FaceNet). Notably, DiffAttack demonstrates superior transferability, surpassing traditional noise-based methods by over 15.28% and semantic-based approaches by approximately 5.21% on benchmark datasets like FFHQ and CelebA-HQ.
| Comments: | Accepted at IEEE International Joint Conference on Biometrics (IJCB) 2026 |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| Cite as: | arXiv:2607.28936 [cs.CV] |
| (or arXiv:2607.28936v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.28936 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Nima Karimian [view email]
[v1]
Fri, 31 Jul 2026 01:33:53 UTC (16,759 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.