UltraVR: A Diagnostic Ultra-Resolution Image–VQA Benchmark for Evidence-Grounded Reasoning
Abstract
Vision-language models (VLMs) have advanced rapidly on visual question answering and multimodal reasoning benchmarks. Yet it remains unclear whether they can reason over ultra-resolution images, where answer-critical evidence may be tiny, subtle, spatially distant, or distributed across regions. Existing evaluations largely report final-answer accuracy, offering limited insight into whether models acquire and integrate the necessary visual evidence. We introduce UltraVR, a diagnostic benchmark for evidence-grounded visual reasoning over ultra-resolution images. UltraVR spans four high-value scenarios: Closed-Circuit Television (CCTV) surveillance, remote sensing (RS), whole-slide image (WSI) pathology, and industrial anomaly detection (AD). These domains pose complementary challenges, including fine-grained object grounding in crowded CCTV scenes, long-range spatial comparison in RS, multi-scale evidence navigation in WSI, and subtle irregularity detection in repetitive industrial layouts. Beyond image-question-answer triples, each UltraVR instance includes a structured ground-truth chain of thought with step-level questions, intermediate answers, and reasoning operation labels. These labels decompose reasoning into evidence grounding, local perception, quantification, evidence integration, and decision inference, enabling process-level diagnosis rather than black-box answer scoring. Using UltraVR, we evaluate frontier proprietary and open-weight VLMs and show that current models remain far from reliable on ultra-resolution reasoning. The structured annotations further localize persistent difficulties in grounding and local perception, even with correct preceding context, while downstream inference often recovers when intermediate visual facts are supplied. These findings establish UltraVR as a diagnostic testbed for measuring not only task success, but where models struggle to establish and use visual evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.