Marginal Utility Token Elimination for Vision Transformers
Abstract
Existing token reduction methods leave their framework choice to ad-hoc heuristics over the backbone's inductive bias or to task-specific learning of an auxiliary reducer module. We instead organize token reduction into two axes, a partition that groups tokens into clusters and an aggregation that produces a representative for each cluster, and state a single explicit objective over both axes: a sensitivity-weighted clustering distortion in the key-projected space, which a first-order Taylor expansion of the next block motivates and three stated approximations make computable. The proposed MUTE minimizes this objective by alternating its two analytic coordinate updates, a weighted-mean aggregation and a nearest-representative partition, which descend the objective monotonically and recover the size bias of bipartite merging as a derived correction rather than as a heuristic. Prior token pruning, merging, and clustering reducers occupy restricted cells of the same design grid, each replacing at least one of the two coordinate updates with a mechanism chosen outside the objective. Under the unified protocol MUTE attains the first-rank training-free classification accuracy, and the same reducer transfers without modification to backbones spanning supervised and self-supervised pretraining and to tasks spanning dense prediction, fine-grained classification, and zero-shot domain shift. A systematic grid search shows that design choices inside the framework make up an equivalent family within multi-seed noise while design choices outside it incur substantial losses, locating the sensitivity of token reduction on the partition axis as the structure of the objective anticipates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.