From Targeted Coverage to Conditional Information for Visual Token Pruning
Abstract
Vision-language models (VLMs) process long sequences of visual tokens, resulting in substantial inference cost. Existing visual token pruning methods typically rely on instruction relevance, visual redundancy, or their combination, but can retain tokens whose evidence relevant to the instruction is already represented by the selected set. We propose TACI (Target-Aware Coverage and Conditional Information), a training-free framework that selects visual evidence according to the additional information it provides about the instruction beyond the tokens already retained. Once the instruction representations have incorporated visual context, TACI evaluates each candidate using conditional mutual information. To avoid carrying the full visual sequence until this point, it first constructs a compact candidate pool using target-weighted visual coverage. Both stages use Gaussian process (GP) models over token representations, providing a unified probabilistic framework for the selection process. Experiments on seven benchmarks with LLaVA and Qwen backbones demonstrate the effectiveness of TACI across different token budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.