Localizing and Disrupting Entity Identification in Vision-Language Models
Abstract
Shown a photograph of a famous person, brand, or landmark, a vision–language model (VLM) can report facts that no pixel contains: it first recognizes the entity and then retrieves what it has memorized about it. The same step lets memory override the image. Which components of a VLM support recognition of familiar entities? We answer this question with two tools: familiarity tiers, derived from the model’s own behavior, which separate entities it knows robustly, knows weakly, or does not know; and an exhaustive sweep of attention-layer knockouts, scored against perceptual control questions on the same images. Across four open VLMs (Qwen2.5-VL, Qwen3-VL, Gemma-4, and LLaVA-1.5), recognition loss exhibits prominent peaks in mid-decoder attention layers, with selectivity relative to perceptual controls varying across models. In Qwen2.5-VL and LLaVA-1.5, where the band is one or two layers wide, ablating it substantially lowers recognition accuracy while having smaller effects on answers to the perceptual control questions: an entity-recognition gate. In Gemma-4 and Qwen3-VL the band spans five layers, and ablating it also changes a substantial share of the control answers. An in-depth analysis of Qwen2.5-VL shows that its gate generalizes across faces, logos, landmarks, and artworks; that ablating it lowers identification of highly-known-tier edited faces from 58% to 15% while reports of the same images’ attributes are unchanged and general VQA accuracy drops by 1–3 points; and that ablating five of its attention heads lowers logo recognition from 96% to 3%. Finally, the gate influences behavior when memory and image disagree: ablating it reduces memory-driven identification errors on VLMBias by up to 24 points, and on people edited to wear a doctor’s coat it shifts the model’s answer from their real occupation toward “doctor”.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.