acceptodds
Under review as a conference paper at ICLR 2027

Correct Yet Insensitive to Vision: Adaptive Intervention to Improve Visual Reliance in VLMs

Abstract

While vision-language models (VLMs) continue to improve on visual question answering (VQA) benchmarks, we identify a critical limitation: an apparently correct answer may not derive from visual evidence. We introduce the Visual Insensitivity Score (VIS) to quantify this behavior and surprisingly find that many correct answers remain unchanged when visual input is replaced or removed. This reveals substantial visual insensitivity in VLMs that conventional VQA accuracy fails to capture. To improve visual reliance, existing methods typically employ visual interventions with manually specified strengths, without explicitly accounting for whether model answers are grounded in visual evidence. To address this limitation, we propose Adaptive Visual Intervention Optimization (VisIO), which explicitly optimizes lightweight interventions to enforce reliance on answer-determining visual evidence rather than other contextual cues. VisIO achieves this by constructing language-preserving, answer-flipping visual counterfactual pairs, where the linguistic context and most visual content remain unchanged while controlled visual changes flip the correct answer. Importantly, VisIO constructs such pairs using a small set of unlabeled target videos and optimizes only hundreds of head-wise parameters before inference. Across multiple VLMs and video VQA benchmarks, VisIO consistently improves accuracy while reducing visual insensitivity and strengthening reliance on visual evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.