Diagnosing Query Awareness in Visual Token Pruning: From Biased Importance Signals to Calibrated Selection
Abstract
Visual-token pruning for VLMs relies on importance scores that are often treated as proxies for question-relevant evidence. Popular scores can be divided into image-conditioned visual priors and query-conditioned signals represented by early text-visual attention. Their high aggregate fidelity has been reported, yet aggregate fidelity alone cannot reveal what these scores measure and limits understanding. In this work, we conduct a structured diagnostic analysis along three axes: prompt sensitivity, selection-behavior preference, and depth-wise query formation. Across multiple VLMs, early text-visual attention exhibits little query discrimination despite being prompt-conditioned in construction. It behaves as an image-conditioned local-detail prior, while visual priors favor broader spatial coverage; the two families succeed on complementary global- and local-evidence regimes. Diversity-aware selection reduces redundancy but does not remove these inherited biases. Reliable query-discriminative attention emerges only in middle-to-late language-model layers. These findings motivate Query-Calibrated Pruning (Q-Cal), a two-stage framework for multi-round VLM inference. Q-Cal first forms a reusable visual prefix from complementary early priors, then refines it with a calibrated mixture of deeper text-visual attention once query-specific evidence has formed. Across six Qwen-family VLMs and eight benchmarks, Q-Cal improves both the accuracy and efficiency of visual pruning for multi-round questions. Code is available at https://github.com/anonymous733882/q-cal-submission.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.