Does Localization Inform Adaptation for Visual Grounding in Vision-Language Models?
Abstract
Correct answers in visual question answering can arise from language priors or irrelevant visual context, so improving visual grounding requires changing how models use evidence. Parameter adaptation offers a route to this change, but leaves open which layers should be updated. Functional localization suggests adapting the layers whose states most strongly influence an answer. This choice assumes that strong state-level influence identifies effective adaptation sites. We test this assumption using a counterfactual VQA dataset that pairs every edit with an operation-matched control and evaluates evidence dependence (Removal), response updates after evidence replacement (Replacement), and invariance to background changes (Context). Across five multimodal large language models (MLLMs; 4B 13B, three architecture families), models respond to relevant evidence but also to irrelevant background changes. Localization and original-state restoration are strongest in early layers, yet similar localization profiles can accompany different single-layer adaptation gains and accuracy costs. On Qwen3.5-9B, localization-ranked placement offers no consistent advantage over random placement at a fixed parameter budget. Joint training with behavior-aligned supervision improves the evaluated grounding metrics, with external abstention gains alongside zero-answer counting and consistency costs. Localization tells us where visual evidence acts on the answer, but not, by itself, where to adapt.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.