Uncertainty-Aware Deepfake Detection via Multi-View Structural Learning
Quick Answer
The proposed uncertainty-aware deepfake detection framework integrates visual, semantic, and structural streams to improve prediction reliability under distribution shifts.
Quick Take
Utilizing Inter-Branch Disagreement Calibration (IBDC), it achieves state-of-the-art generalization on multiple out-of-distribution benchmarks, enhancing calibration and selective prediction performance, particularly for security-critical applications.
Key Points
- Framework combines visual, semantic, and structural evidence streams for deepfake detection.
- Introduces Inter-Branch Disagreement Calibration (IBDC) for uncertainty modeling.
- Achieves state-of-the-art performance on out-of-distribution benchmarks.
- Improves calibration and selective prediction for security-critical applications.
- Demonstrated effectiveness through extensive cross-dataset experiments.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Security-critical biometric and forensic applications require accurate predictions and reliable confidence estimates, particularly under distribution shift. This challenge is especially acute for deepfake detection, where foundation-model-based detectors often exhibit overconfident predictions on out-of-distribution manipulations, which limits their suitability for operational deployment. We propose an uncertainty-aware deepfake detection framework that identifies manipulations through inconsistencies across complementary evidence sources. The framework integrates three streams: a visual stream based on an adapted CLIP encoder, a semantic stream that models consistency among facial attributes through differentiable constraints, and a structural stream that captures class-dependent dependency patterns between semantic and forensic features. To effectively combine these signals, we introduce Inter-Branch Disagreement Calibration (IBDC), a disagreement-aware uncertainty modeling mechanism that links predictive uncertainty to conflicts among evidence streams. Extensive cross-dataset experiments using FaceForensics++ as the training source demonstrate that the proposed framework achieves state-of-the-art generalization across multiple out-of-distribution benchmarks while consistently improving calibration and selective prediction performance. These results show that combining complementary evidence with disagreement-aware uncertainty provides a robust foundation for trustworthy and well-calibrated deepfake detection under distribution shift.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.28769 [cs.CV] |
| (or arXiv:2607.28769v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.28769 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Muhammad Umar Farooq [view email]
[v1]
Thu, 30 Jul 2026 18:42:01 UTC (4,808 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.