UltraVR: A Diagnostic Ultra-Resolution Image-VQA Benchmark for Evidence-Grounded Reasoning
Quick Answer
UltraVR introduces a benchmark for evaluating vision-language models (VLMs) on ultra-resolution images, revealing significant shortcomings in evidence-grounded reasoning.
Quick Take
Current models struggle with tasks like fine-grained object grounding and spatial comparisons, indicating a need for improved visual evidence integration. This benchmark allows for detailed diagnostics of model failures, particularly in evidence grounding and local perception.
Key Points
- UltraVR benchmarks across four scenarios: CCTV, remote sensing, pathology, and anomaly detection.
- Structured annotations in UltraVR enable detailed process-level diagnostics of reasoning failures.
- Current VLMs show unreliable performance on ultra-resolution reasoning tasks.
- Errors are primarily found in evidence grounding and local perception stages.
- Downstream inference often improves when intermediate visual facts are provided.
Paper Resources
📖 Reader Mode
~2 min readAbstract:Vision-language models (VLMs) excel on visual question answering and multimodal reasoning benchmarks. Yet their capability on ultra-resolution images - where critical evidence is tiny, subtle, spatially distant, or distributed - remains unclear. Existing evaluations largely report final-answer accuracy, offering limited insight into whether models acquire and integrate the necessary visual evidence. We introduce UltraVR, a diagnostic benchmark for evidence-grounded visual reasoning over ultra-resolution images. UltraVR spans four high-value scenarios: CCTV surveillance, remote sensing (RS), whole-slide image (WSI) pathology, and industrial anomaly detection (AD). These domains pose complementary challenges: fine-grained object grounding in crowded CCTV scenes, long-range spatial comparison in RS, multi-scale evidence navigation in WSI, and subtle irregularity detection in repetitive industrial layouts. Beyond standard QA triples, each instance includes a structured ground-truth chain of thought with step-level questions, intermediate answers, and reasoning labels. These labels decompose reasoning into evidence grounding, local perception, quantification, evidence integration, and decision inference, enabling process-level diagnosis over black-box scoring. Using UltraVR, we evaluate frontier VLMs and show that current models remain far from reliable on ultra-resolution reasoning. Importantly, the structured annotations allow us to localize failures across the visual-to-decision pipeline: errors concentrate in evidence grounding and local perception, while downstream inference often recovers when intermediate visual facts are supplied. These findings demonstrate UltraVR as a diagnostic testbed for measuring not only whether VLMs answer correctly, but where their ultra-resolution reasoning process breaks.
| Comments: | 10 pages, 1 figure |
| Subjects: | Computer Vision and Pattern Recognition (cs.CV) |
| Cite as: | arXiv:2606.05576 [cs.CV] |
| (or arXiv:2606.05576v1 [cs.CV] for this version) | |
| https://doi.org/10.48550/arXiv.2606.05576 arXiv-issued DOI via DataCite |
Submission history
From: Gexin Huang [view email]
[v1]
Thu, 4 Jun 2026 01:51:16 UTC (4,833 KB)
— Originally published at arxiv.org
Want this in your inbox every morning?
Daily brief at your local 8am — bilingual EN/中文, free.
More from arXiv cs.CV
See more →ProMoE-FL: Prototype-conditioned Mixture of Experts for Multimodal Federated Learning with Missing Modalities
ProMoE-FL introduces a Prototype-conditioned Mixture-of-Experts framework for multimodal federated learning, effectively addressing missing modalities. It outperforms existing methods on four chest X-ray datasets, demonstrating superior feature synthesis capabilities in both homogeneous and heterogeneous settings.