Diff-ID: Identity Consistent Facial Image Generation and Morphing via Diffusion Models
Quick Answer
Diff-ID is a novel diffusion-based framework for generating high-resolution facial images that ensures identity consistency, utilizing a custom 210K image dataset and integrating ArcFace and CLIP embeddings.
Quick Take
While it does not surpass InstantID in raw ArcFace Face Similarity, it achieves lower FID and superior identity-realism trade-offs, emphasizing the need for joint evaluation of identity preservation and photorealism.
Key Points
- Diff-ID uses a custom dataset of 210K images from CelebA-HQ, FFHQ, and LAION-Face.
- Integrates ArcFace and CLIP embeddings with a dual cross attention adapter.
- Achieves lower FID scores while maintaining identity realism compared to InstantID.
- Introduces a pseudo discriminator loss based on ArcFace cosine similarity.
- Proposes a DDIM-based morphing pipeline for qualitative facial interpolation.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Generative diffusion models have revolutionized facial image synthesis, yet robust identity preservation in high resolution outputs remains a critical challenge. This issue is especially vital for security systems, biometric authentication, and privacy sensitive applications, where any drift in identity integrity can undermine trust and functionality. We introduce Diff-ID, a diffusion based framework that enforces identity consistency while delivering photorealistic quality. Central to our approach is a custom 210K image dataset synthesized from CelebA-HQ, FFHQ, and LAION-Face and captioned via a fine tuned BLIP model to bolster identity awareness during training. Diff-ID integrates ArcFace and CLIP embeddings through a dual cross attention adapter within a fine tuned Stable Diffusion UNet. To further reinforce identity fidelity, we propose a pseudo discriminator loss based on ArcFace cosine similarity with exponential timestep weighting. Experiments on held out and unseen faces show that Diff-ID does not exceed InstantID in raw ArcFace Face Similarity, but achieves substantially lower FID and the strongest FIQ based identity--realism trade off among the evaluated methods. We also present a unified DDIM based morphing pipeline that enables qualitative facial interpolation without per identity fine tuning. We further argue that identity preservation and photorealism should be evaluated jointly rather than in isolation, as high identity similarity alone does not guarantee realistic outputs. To make this trade off explicit, we report Face Image Quality (FIQ) as a complementary ratio based score that combines identity similarity and perceptual realism while keeping FS and FID as the primary metrics.
| Comments: | 20 pages, 11 figures, 3 tables |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.25078 [cs.CV] |
| (or arXiv:2607.25078v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.25078 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Taimoor Rizwan [view email]
[v1]
Mon, 27 Jul 2026 21:13:29 UTC (27,098 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.