acceptodds
Under review as a conference paper at ICLR 2027

Your Large Vision-Language Model Already Knows How to Localize Changes

Abstract

Visual change understanding necessitates both seeing and saying. However, existing approaches typically realize only one of these capabilities: LVLMs excel at saying changes but struggle to precisely see them, whereas localization-only methods can accurately see changes but lack the semantic capacity to say what those changes represent. To fill this gap, we endow LVLMs with the ability to reliably see changes. By investigating the attention head representations in frozen LVLMs, we find that a small subset of attention heads ( of all attention heads) already encodes change-informative spatial evidence. We excavate these heads by measuring their attention sum deviation from an unchanged self-reference baseline and their spatial consistency across both image directions. Despite being selected without pixel-level supervision, these heads produce responses that closely align with the ground truth. Furthermore, we diffuse their attention through semantic affinities extracted from the frozen vision encoder—rendering stable maps for effective SAM refinements. Extensive experiments across diverse change localization benchmarks demonstrate that our training-free framework substantially surpasses LVLM-based baselines and matches dedicated localization-only approaches. Our results suggest that LVLMs encode an innate spatial understanding of changes via distinct attention patterns, which, combined with their world knowledge and strong descriptive capabilities, enables comprehensive and generalizable change understanding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.