Correct Yet Insensitive to Vision: Adaptive Intervention to Improve Visual Reliance in VLMs
Abstract
While vision-language models (VLMs) continue to improve on visual question answering (VQA) benchmarks, we identify a critical limitation: an apparently correct answer may not derive from visual evidence. We introduce the Visual Insensitivity Score (VIS) to quantify this behavior and surprisingly find that many correct answers remain unchanged when visual input is replaced or removed. This reveals substantial visual insensitivity in VLMs that conventional VQA accuracy fails to capture. To improve visual reliance, existing methods typically employ visual interventions with manually specified strengths, without explicitly accounting for whether model answers are grounded in visual evidence. To address this limitation, we propose Adaptive Visual Intervention Optimization (VisIO), which explicitly optimizes lightweight interventions to enforce reliance on answer-determining visual evidence rather than other contextual cues. VisIO achieves this by constructing language-preserving, answer-flipping visual counterfactual pairs, where the linguistic context and most visual content remain unchanged while controlled visual changes flip the correct answer. Importantly, VisIO constructs such pairs using a small set of unlabeled target videos and optimizes only hundreds of head-wise parameters before inference. Across multiple VLMs and video VQA benchmarks, VisIO consistently improves accuracy while reducing visual insensitivity and strengthening reliance on visual evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.