CentroidPrune: Centroid Matching for Visual Token Pruning in Vision-Language Models
Abstract
Vision-Language Models (VLMs) incur substantial inference cost because a single image is converted into a large number of visual tokens. Token pruning addresses this by removing less informative visual tokens. Existing methods rank tokens with an importance measure such as attention score or token similarity, and retain the top-ranked ones. However, these measures do not reflect how removing tokens changes the model's output, and existing methods degrade sharply as the pruning ratio increases. In this work, we analyze the effect of pruning on the attention output and show that the induced change decomposes into two factors: the removed attention mass and the directional deviation of the removed tokens. Attention-based pruning minimizes the former while leaving the latter uncontrolled, and therefore does not minimize the change in the attention output. Motivated by this analysis, we propose CentroidPrune, a training-free visual token pruning method. We formulate token pruning as a subset selection problem that directly minimizes the distance between the original attention output and the output after pruning and renormalization. Extensive experiments on various VLMs and multimodal benchmarks show that CentroidPrune consistently outperforms prior methods, retaining 98.4% of the unpruned performance while discarding up to 94.4% of the visual tokens.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.