acceptodds
Under review as a conference paper at ICLR 2027

UltraVR: A Diagnostic Ultra-Resolution Image–VQA Benchmark for Evidence-Grounded Reasoning

Abstract

Vision-language models (VLMs) have advanced rapidly on visual question answering and multimodal reasoning benchmarks. Yet it remains unclear whether they can reason over ultra-resolution images, where answer-critical evidence may be tiny, subtle, spatially distant, or distributed across regions. Existing evaluations largely report final-answer accuracy, offering limited insight into whether models acquire and integrate the necessary visual evidence. We introduce UltraVR, a diagnostic benchmark for evidence-grounded visual reasoning over ultra-resolution images. UltraVR spans four high-value scenarios: Closed-Circuit Television (CCTV) surveillance, remote sensing (RS), whole-slide image (WSI) pathology, and industrial anomaly detection (AD). These domains pose complementary challenges, including fine-grained object grounding in crowded CCTV scenes, long-range spatial comparison in RS, multi-scale evidence navigation in WSI, and subtle irregularity detection in repetitive industrial layouts. Beyond image-question-answer triples, each UltraVR instance includes a structured ground-truth chain of thought with step-level questions, intermediate answers, and reasoning operation labels. These labels decompose reasoning into evidence grounding, local perception, quantification, evidence integration, and decision inference, enabling process-level diagnosis rather than black-box answer scoring. Using UltraVR, we evaluate frontier proprietary and open-weight VLMs and show that current models remain far from reliable on ultra-resolution reasoning. The structured annotations further localize persistent difficulties in grounding and local perception, even with correct preceding context, while downstream inference often recovers when intermediate visual facts are supplied. These findings establish UltraVR as a diagnostic testbed for measuring not only task success, but where models struggle to establish and use visual evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.