A Step Forward Towards Trustworthy Risk-Aware Facial Retrieval (RA-FR)
Quick Answer
The proposed Risk-Aware Facial Retrieval (RA-FR) framework enhances facial image retrieval in surveillance by ensuring ground truth inclusion within specified risk levels.
Quick Take
It integrates blind face restoration, robust feature extraction using DINOv1 ViT-B, and dynamic calibration through conformal prediction, achieving a consistent 5% risk target on the IMFDB benchmark with an average retrieval set size of 10 images.
Key Points
- RA-FR guarantees ground truth inclusion with user-defined risk and confidence levels.
- It reduces aleatoric uncertainty using a hybrid blind face restoration technique.
- Features are extracted using self-supervised DINOv1 ViT-B with GGeM pooling.
- Dynamic calibration of retrieval set sizes is achieved through conformal prediction.
- The framework consistently meets a 5% risk target on the IMFDB benchmark.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Facial image retrieval in unconstrained surveillance environments is a high-stakes challenge where missing a subject of interest -- a single false negative -- is simply not an option. Despite near-perfect performance on curated benchmarks, current recognition systems falter under real-world domain shifts such as low resolution, motion blur, and uncontrolled illumination (e.g., SCFace). Addressing this reliability gap, we propose Risk-Aware Facial Retrieval (RA-FR), a framework that moves beyond fixed Top-$k$ retrieval to adaptive set generation, guaranteeing ground truth inclusion within a user-specified risk level ($\alpha$) and confidence level ($1 - \delta$). Our approach integrates three core contributions: (1) reducing aleatoric uncertainty via a hybrid blind face restoration technique coupling Latent Consistency Models (InterLCM) and DiffBIR; (2) extracting discriminative, restoration-robust features via self-supervised DINOv1 ViT-B with GGeM pooling; and (3) employing conformal prediction with Hoeffding's inequality to dynamically calibrate retrieval set sizes based on query uncertainty. On the IMFDB benchmark, it consistently satisfies a 5% risk target with an average retrieval set size of approximately 10 images. By unifying domain-specific restoration, robust representation learning, and provable decision rules, RA-FR offers a pipeline that makes facial retrieval in surveillance both reliable and auditable. The code is available at: this https URL.
| Comments: | Accepted at the 28th International Conference on Pattern Recognition (ICPR 2026). Extended version with additional citations and expanded sections |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| ACM classes: | I.2.10; I.4.8; I.4.9 |
| Cite as: | arXiv:2607.16279 [cs.CV] |
| (or arXiv:2607.16279v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.16279 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Muhammad Emmad Siddiqui [view email]
[v1]
Thu, 9 Jul 2026 14:55:45 UTC (16,702 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.