Change the Image, Change the Answer? Predicting VLM Responses to Regional Edits
Abstract
A vision–language model that answers an image question correctly may follow a regional edit intended to change the answer-relevant facts, or retain its original answer. We ask whether original-image measurements can rank these responses in multiple-choice VQA without observing the edited image. We propose GLANCE, a three-view score fixed before evaluation contrasting original-answer support on the full image and an enlarged regional crop with support remaining after masking. Edited images supply labels and determine admission to the primary cohorts but never enter the score. On six model-specific Visual7W cohorts selected by post-edit checks, GLANCE ranks target-following above preservation with AUROC 0.707–0.832 and exceeds the strongest tested attention and geometry readouts in five. On two models, GLANCE also ranks changes among originally correct answers without applying crop or placebo admission to the new edits. In matched diagnostic probes, adding masked support to visible supports raises AUROC by 0.124–0.243 across six Visual7W and two single-layer MEGA cohorts, with all refitting-sensitivity intervals above zero. Positive increments also persist in two small subsets where a majority of human raters selects the target answer. Component analyses show that masked-view access matters more than the exact three-view weighting; the tested binary answer-agreement rules generally retain less ranking information than continuous support. The score ranks response tendencies under the tested edits; among changed answers, it does not consistently separate target from third-option responses.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.