acceptodds
Under review as a conference paper at ICLR 2027

CentroidPrune: Centroid Matching for Visual Token Pruning in Vision-Language Models

Abstract

Vision-Language Models (VLMs) incur substantial inference cost because a single image is converted into a large number of visual tokens. Token pruning addresses this by removing less informative visual tokens. Existing methods rank tokens with an importance measure such as attention score or token similarity, and retain the top-ranked ones. However, these measures do not reflect how removing tokens changes the model's output, and existing methods degrade sharply as the pruning ratio increases. In this work, we analyze the effect of pruning on the attention output and show that the induced change decomposes into two factors: the removed attention mass and the directional deviation of the removed tokens. Attention-based pruning minimizes the former while leaving the latter uncontrolled, and therefore does not minimize the change in the attention output. Motivated by this analysis, we propose CentroidPrune, a training-free visual token pruning method. We formulate token pruning as a subset selection problem that directly minimizes the distance between the original attention output and the output after pruning and renormalization. Extensive experiments on various VLMs and multimodal benchmarks show that CentroidPrune consistently outperforms prior methods, retaining 98.4% of the unpruned performance while discarding up to 94.4% of the visual tokens.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.