acceptodds
Under review as a conference paper at ICLR 2027

VIGOR: Visual Intervention-Grounded Evaluation for Observation and Reasoning

Abstract

Visual reasoning benchmarks are read as evidence that a vision-language model (VLM) answers from the image, yet final-answer accuracy cannot show whether the image was used. Intervening on the image and reading the change in accuracy is the usual remedy, but an accuracy change can reflect a shift in the model's response prior as well as the removal of task-relevant content. We ask when such a delta supports a claim that the intended evidence was used. We introduce VIGOR, an audit that evaluates each item under five views and summarizes them as a validity profile, and VIGOR-Stress300, an evidence-annotated subset of MMStar, BLINK-Twice, and MMPerspective. Across open VLMs from three families, masking the annotated evidence lowers accuracy on MMStar and raises it on BLINK-Twice; read on its own, the gain would be taken as improved use of the image. Two requirements make the deltas interpretable. Under score legibility, a panel of option structure, label marginal, and answers without evidence shows that the BLINK-Twice gain is a response shift toward what the models answer for an occluded image, that a text-only polarity rule outscores every masked view, and that original-view predictions show no positive chance-adjusted agreement with the labels. Under matched intervention, MMStar's degradation survives label adjustment, six removal styles, location- and magnification-matched controls, and boxes drawn by independent auditors. On MMPerspective the tested interventions leave a small, model-dependent effect unresolved. We recommend that intervention-based audits report the legibility panel and the matched controls with every delta.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.