Attending Without Selecting: Visual Key Collapse and Hallucination in Large Vision-Language Models
Abstract
Large vision-language models often describe objects that are absent from the image, a failure usually attributed to insufficient attention to visual input. We find that hallucinated and faithful captions place the same share of attention on the image, but in hallucinated captions that attention depends less on the word being generated. The failure lies in selecting where to look rather than in how much to look. We trace it to a collapse of visual keys onto a few dominant directions, which leads every query to rank image regions by nearly the same criteria. The collapse appears across architectures with different visual interfaces, does not depend on the relative norms of visual and text keys, and is more severe for images that the model describes less faithfully. Building on this diagnosis, we introduce Whitened Attention Steering, a training-free method that suppresses the dominant directions in visual attention scores while keeping the total attention on the image fixed. Hallucination decreases steadily as the correction is strengthened and increases when it is reversed, whereas matched corrections of other directions or other heads have little effect. On LLaVA-1.5, the share of captions containing a hallucinated object falls from 54% to 30%, and across models the gain follows the severity of collapse. These findings indicate that hallucination in current vision-language models stems less from a shortage of visual attention than from a loss of selection within it.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.