acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing Query Awareness in Visual Token Pruning: From Biased Importance Signals to Calibrated Selection

Abstract

Visual-token pruning for VLMs relies on importance scores that are often treated as proxies for question-relevant evidence. Popular scores can be divided into image-conditioned visual priors and query-conditioned signals represented by early text-visual attention. Their high aggregate fidelity has been reported, yet aggregate fidelity alone cannot reveal what these scores measure and limits understanding. In this work, we conduct a structured diagnostic analysis along three axes: prompt sensitivity, selection-behavior preference, and depth-wise query formation. Across multiple VLMs, early text-visual attention exhibits little query discrimination despite being prompt-conditioned in construction. It behaves as an image-conditioned local-detail prior, while visual priors favor broader spatial coverage; the two families succeed on complementary global- and local-evidence regimes. Diversity-aware selection reduces redundancy but does not remove these inherited biases. Reliable query-discriminative attention emerges only in middle-to-late language-model layers. These findings motivate Query-Calibrated Pruning (Q-Cal), a two-stage framework for multi-round VLM inference. Q-Cal first forms a reusable visual prefix from complementary early priors, then refines it with a calibrated mixture of deeper text-visual attention once query-specific evidence has formed. Across six Qwen-family VLMs and eight benchmarks, Q-Cal improves both the accuracy and efficiency of visual pruning for multi-round questions. Code is available at https://github.com/anonymous733882/q-cal-submission.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.