acceptodds
Under review as a conference paper at ICLR 2027

Vision-Language Models Do Not Keep Looking

Abstract

Autoregressive Large Vision-Language Models solve multimodal reasoning tasks by retaining hundreds or thousands of visual tokens, allowing every generated token to attend to the visual tokens at any time during the generation process. However, whether continuous access to these visual tokens is necessary remains poorly understood. In this work, we investigate how visual token influence evolves along two dimensions: generation time and transformer depth. Using long-form visual reasoning benchmarks and multiple model scales, we first analyze the attention dynamics across generation time and transformer depth. We identify a temporal and layer-wise structure in the use of visual information. Visual attention peaks during the initial generation steps and then rapidly declines as attention shifts toward previously generated reasoning tokens. Across depth, visual influence is concentrated in intermediate layers and becomes very weak in the final layers. Motivated by these observations, we propose a simple intervention that removes all visual tokens across generation time and model depth. These interventions confirm a transition from early multimodal integration, during which visual information is incorporated into the generated token representations, to later generation, where the model relies primarily on these accumulated representations rather than direct access to the visual tokens. Thus, retaining the complete visual context at every layer and decoding step may be unnecessary. Our experiments demonstrate that we can remove up to 59% of visual tokens across generation time and 11% across model depth, while incurring only a negligible performance drop.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.