acceptodds
Under review as a conference paper at ICLR 2027

Now You See It, Now You Don’t: Rethinking visual token allocation across Vision Language Model layers

Abstract

Vision-language models (VLMs) can often localize relevant regions in an image, yet still fail to perceive the fine-grained visual evidence needed to answer correctly, leading to visual hallucinations. Using more tokens provides a finer representation of the image, but also increases the computation, memory, and latency required for inference. VLMs therefore operate under a limited visual-token budget: spending more of it on irrelevant details wastes computation without improving the answer. The central challenge is therefore to spend this limited budget where additional visual detail is most useful. A common strategy is attention-guided cropping, however, existing approaches typically extract this attention from a manually selected layer or a fixed range of layers, without establishing which depth provides the most reliable signal for visual selection. We show that the choice of layer is a dominant factor in attention-guided visual selection. By systematically masking text-to-image attention across network depth, we show that the information relevant to the model's prediction is concentrated within an early-to-middle portion of the network. This boundary occurs substantially earlier than the layers used by existing attention-based selection methods. Based on this finding, we introduce two complementary strategies. Depth-Weighted Attention (DWA) combines attention information across the full depth of the model through a closed-form ridge regression, learning layer weights that suppress less informative late-layer contributions and produce more reliable spatial selection. Adaptive Visual Resolution (AVR) uses the identified depth boundary to remove visual tokens that no longer affect the prediction and redirects the saved computation toward higher-resolution encoding, allowing the model to represent important image regions with finer visual detail. Across nine matched-compute comparisons, DWA consistently improves over existing layer-selection strategies, achieving gains of 7.9 and 8.9 percentage points over the strongest baseline and locating answer-relevant evidence in 63.9% of cases compared with 13.6%. Against native dynamic-resolution models operating at 5.5 the visual-token budget, DWA and AVR achieve statistically indistinguishable performance while using only 14-18% of the visual-token computation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.