Vision-Language Models Do Not Keep Looking
Abstract
Autoregressive Large Vision-Language Models solve multimodal reasoning tasks by retaining hundreds or thousands of visual tokens, allowing every generated token to attend to the visual tokens at any time during the generation process. However, whether continuous access to these visual tokens is necessary remains poorly understood. In this work, we investigate how visual token influence evolves along two dimensions: generation time and transformer depth. Using long-form visual reasoning benchmarks and multiple model scales, we first analyze the attention dynamics across generation time and transformer depth. We identify a temporal and layer-wise structure in the use of visual information. Visual attention peaks during the initial generation steps and then rapidly declines as attention shifts toward previously generated reasoning tokens. Across depth, visual influence is concentrated in intermediate layers and becomes very weak in the final layers. Motivated by these observations, we propose a simple intervention that removes all visual tokens across generation time and model depth. These interventions confirm a transition from early multimodal integration, during which visual information is incorporated into the generated token representations, to later generation, where the model relies primarily on these accumulated representations rather than direct access to the visual tokens. Thus, retaining the complete visual context at every layer and decoding step may be unnecessary. Our experiments demonstrate that we can remove up to 59% of visual tokens across generation time and 11% across model depth, while incurring only a negligible performance drop.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.