When Pruning Rankings Reverse: State-Conditioned Visual Token Compression across Architectures
Abstract
The high cost of visual tokens makes token compression key to accelerating VLM inference. Existing methods mainly select tokens by token-wise task relevance or inter-token diversity, and typically apply a fixed rule to all inputs. However, we find that the relative effectiveness of these two representative approaches reverses across model architectures and shifts with task and compression ratio; even under the same architecture and task, preferences differ across samples. This indicates that token selection should adapt to each sample rather than follow a fixed rule. We therefore propose SATOR (Sample-Adaptive Token Reduction), a training-free, sample-adaptive token selection method that requires no architecture- or task-specific parameters. SATOR estimates the task contribution of each token from attention and Value norms. Before selection, it uses the visual redundancy and task support of the current input to balance two marginal gains: a distinctiveness gain that suppresses redundancy with retained tokens, and a representativeness gain that encourages coverage of unretained task-relevant information. On LLaVA-NeXT-7B, SATOR retains only 11.1% of visual tokens while preserving 97.5% of the original average performance, achieving a 3.24 prefill speedup, a 2.19 end-to-end speedup, and a 6.02 reduction in KV-cache size.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.