SLAPBench: Benchmarking Multimodal Large Language Models for Four-Finger SLAP Fingerprint Verification
Quick Answer
SLAPBench introduces the first benchmark for verifying four-finger SLAP fingerprints using multimodal large language models (MLLMs).
Quick Take
Evaluations show that Claude Opus 4.8 achieves the best binary result with a 20.2% False Accept Rate, while Qwen3-VL-8B attains perfect separation (AUC = 1.000). The study highlights the impact of prompting on model performance and fairness disparities across demographics.
Key Points
- SLAPBench is built from NIST SD302b with 7,832 fingerprint pairs.
- Task-description prompting leads to nearly 100% False Accept Rate in open-source models.
- Claude Opus 4.8 shows the best performance with a 20.2% FAR.
- Qwen3-VL-8B achieves perfect separation (AUC = 1.000) but is diagnostic.
- Fairness analysis reveals growing disparities as discrimination weakens.
DeepSignal Analysis
What happened
SLAPBench is introduced as the first benchmark for verifying four-finger SLAP fingerprints using multimodal large language models (MLLMs). Evaluations reveal that Claude Opus 4.8 achieves a 20.2% False Accept Rate, while Qwen3-VL-8B shows perfect separation with an AUC of 1.000. The study emphasizes the role of prompting in model performance and fairness disparities.
Key evidence
- SLAPBench evaluates four open-source MLLMs and Claude Opus 4.8 using a dataset of 7,832 fingerprint pairs, consisting of 176 mated and 7,656 non-mated pairs.
- Under task-description prompting, all four open-source models exhibit near-100% False Accept Rate, while Claude Opus 4.8 maintains a 20.2% False Accept Rate.
- Qwen3-VL-8B achieves an AUC of 1.000, indicating perfect separation, but this is attributed to the dataset's structure rather than inherent model capability.
Why it matters
The introduction of SLAPBench provides a foundational benchmark for assessing MLLMs in the context of fingerprint verification, an area previously unexplored. The findings highlight significant performance variations among models, particularly under different prompting strategies. Understanding these disparities is crucial for developing fair and effective identity verification systems, especially in sensitive applications like border control and law enforcement.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Four-finger SLAP fingerprints are flat live-scan impressions of the index, middle, ring, and little fingers of one hand, used for identity verification in border control and law enforcement. No benchmark has evaluated whether multimodal large language models (MLLMs) can verify identity from SLAP images. We introduce SLAPBench, the first benchmark for MLLM-based four-finger SLAP fingerprint verification, built from NIST SD302b with 7,832 pairs (176 mated, 7,656 non-mated). We evaluate four open-source MLLMs (InternVL3-8B, Qwen2.5-VL-7B, Qwen3-VL-8B, Gemma-3-12B) and the proprietary Claude Opus 4.8 under zero-shot, task-description, and similarity-scoring prompts.
Prompting governs verification behavior. Task-description prompting collapses all four open-source models to near-100% False Accept Rate (FAR), and Gemma-3-12B collapses under zero-shot as well; Claude Opus 4.8 alone resists collapse under both binary prompts, giving the best binary result (FAR = 20.2%). Similarity scoring removes collapse across the open-source models and exposes wide capability gaps: Claude reaches AUC = 0.953 and Gemma-3-12B 0.837, while InternVL3-8B is inverted (AUC = 0.590) and Qwen2.5-VL-7B near random (0.567). Qwen3-VL-8B attains perfect separation (AUC = 1.000), which we treat as a diagnostic rather than as capability: SD302b holds one SLAP capture per finger position, so mated pairs are cross-resolution. A matched-resolution control leaves the perfect score intact, ruling out the resolution shortcut; what cannot be excluded within SD302b is near-duplicate detection, since a mated pair is one capture rendered twice. A fairness probe over gender, race, and age suggests disparity grows as discrimination weakens. SLAPBench establishes the first SLAP-specific MLLM baseline and shows that prompting governs collapse while model capability governs discrimination.
| Comments: | 19 pages, 6 figures, 2 tables. Includes appendix with supporting figures and per-subgroup fairness detail. Code and data: this https URL |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV); Artificial Intelligence (cs.AI) |
| ACM classes: | I.4.9; I.5.4; K.6.5 |
| Cite as: | arXiv:2607.15517 [cs.CV] |
| (or arXiv:2607.15517v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.15517 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Bibesh Pyakurel [view email]
[v1]
Fri, 17 Jul 2026 00:04:53 UTC (609 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.