PerceptionBench: Evaluating Atomic Visual Perception in Multimodal Large Language Models
Quick Answer
PerceptionBench introduces a novel benchmark to evaluate atomic visual perception in Multimodal Large Language Models (MLLMs), revealing that no model achieves over 60% accuracy in isolated perceptual tasks.
Quick Take
The study identifies ten atomic perceptual capabilities and constructs 3,000 targeted questions, highlighting significant gaps in current MLLM performance, particularly in perception-related hallucinations.
Key Points
- PerceptionBench evaluates atomic visual perception capabilities of MLLMs across 42 benchmarks.
- Constructed an error taxonomy defining ten atomic perceptual capabilities.
- 3,000 verified questions isolate single capabilities, focusing on perception over reasoning.
- No MLLM reached 60% accuracy, with perception-related hallucination being the weakest area.
- Provides a capability-level standard for diagnosing visual perception in MLLMs.
DeepSignal Analysis
What happened
PerceptionBench is a new benchmark aimed at evaluating atomic visual perception in Multimodal Large Language Models (MLLMs). The study found that no MLLM achieved over 60% accuracy in isolated perceptual tasks, highlighting significant performance gaps, especially in perception-related hallucinations.
Key evidence
- PerceptionBench identifies ten atomic perceptual capabilities and constructs 3,000 targeted questions to evaluate these capabilities.
- The benchmark results show that no model reaches 60% accuracy, indicating that atomic perception remains largely unresolved.
- Perception-related hallucination is identified as the weakest capability on average among the evaluated MLLMs.
Why it matters
The introduction of PerceptionBench addresses the limitations of existing benchmarks that conflate perceptual errors with reasoning failures. By isolating perceptual capabilities, it provides a clearer understanding of MLLMs' limitations in visual perception, which is crucial for advancing the field.
What to watch
Paper Resources
📖 Reader Mode
~2 min readAuthors:Zichao Lin, Yifeng Xie, Bowen Qu, Haiming Wang, Jia Li, Haoning Wu, Yuhao Dong, Zuhao Yang, Jinguo Zhu, Haoyu Lu, Zijia Zhao, Tongtian Yue, Zhangyang Qi, Junwei Yang, Mengfan Dong, Peizhou Cao, Chenzhuang Du, Zaida Zhou, Haotian Yao, Hao Yang, Hongcheng Gao, Lin Sui, Weihong Li, Xinxing Zu, Jia Chen, Yao Wang, Xiaoxue Wu, Yalin Wang, Y. Charles, Yiping Bao, Yangyang Liu, Zhiqi Huang, Xinyu Zhou
Abstract:We introduce PerceptionBench, a benchmark specifically designed to evaluate the atomic visual perception capabilities of Multimodal Large Language Models (MLLMs). Existing benchmarks often fail to isolate perception: holistic evaluations conflate perceptual errors with failures in reasoning or domain knowledge, while application-driven benchmarks only cover narrow, fragmented domains shaped by heuristic designs. To address these limitations, PerceptionBench adopts a bottom-up approach: by diagnosing the earliest failure points in the responses of frontier MLLMs across 42 existing benchmarks, we construct an error taxonomy whose perception branch defines ten atomic perceptual capabilities. Guided by this taxonomy, we construct 3,000 verified questions with short, unambiguous answers, each isolating a single capability, with difficulty stemming from perception rather than reasoning or knowledge. Benchmark results across sixteen frontier MLLMs reveal that atomic perception remains largely unsolved---no model reaches 60\% accuracy, perception-related hallucination is the weakest capability on average, and similar overall scores conceal sharply divergent capability profiles. PerceptionBench thus provides a capability-level standard for measuring and diagnosing the visual perception boundaries of MLLMs.
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2607.24957 [cs.CV] |
| (or arXiv:2607.24957v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2607.24957 arXiv-issued DOI via DataCite (pending registration) |
Submission history
From: Bowen Qu [view email]
[v1]
Mon, 27 Jul 2026 18:04:54 UTC (10,386 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.