acceptodds
Under review as a conference paper at ICLR 2027

Do Vision-Language Models Use the Regions They Cite? A Step-Specific Intervention Test

Abstract

Vision-language models increasingly attach a bounding box to each reasoning step as the visual evidence for that step. Localization metrics measure whether such boxes match reference regions; they do not test whether the step is selectively sensitive to the region it cites. We introduce a step-specific intervention test that requires no region annotations: perturb each cited region in turn, record the change in every fixed step's mean per-token log-likelihood, and compare a step's sensitivity to its own region with its sensitivity to distinct regions cited by other steps of the same chain. Because every pair of distinct regions is compared in both directions, the expected own-region ordering probability is at most 50% under an additive step-and-region sensitivity model with row-exchangeable residuals, so a confidence interval entirely above 50% is evidence against that joint null. For Qwen3-VL-4B the own-region ordering probability is about 61% in each of three image-disjoint GQA cohorts and 72% on A-OKVQA. On the largest GQA cohort, every preregistered change of scoring context and intervention operator (removing the preceding steps, replacing blur with within-box patch permutation, or both) keeps the interval above 50% (lowest point estimate 58%). On a held-out cohort, under a protocol specified in advance, we regenerate each current step with the preceding steps held fixed. The fixed-text margin predicts the relative content change caused by blurring the step's own rather than a sibling step's region (Spearman's ρ = 0.504, 95% CI [0.402, 0.600]). The advantage is strong for GLM-4.6V-Flash on two image-disjoint cohorts, weak and cohort-variable for InternVL3.5-4B, and partial throughout: about 40% of Qwen3-VL-4B's comparisons in the three fixed-text GQA cohorts favor another step's region. A cited region is a testable claim about evidence use, not proof of it.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.