acceptodds
Under review as a conference paper at ICLR 2027

Beyond Token-wise Importance: Vision Token Pruning with Local Relational Redundancy

Abstract

Multimodal Large Language Models (MLLMs) process hundreds or even thousands of visual tokens, resulting in substantial computational and memory overhead. Since these tokens contribute unequally to response generation, *vision token pruning* aims to reduce this overhead by retaining only the most informative tokens. Existing methods typically estimate token importance from individual attention scores, visual features, or token-level redundancy. In this paper, we propose worst-case local dissimilarity as a simple prior for token selection, where high similarity even to a token’s -th most similar visual token indicates strong redundancy. Based on this prior, we introduce Pruner, a training-free vision token pruning method that uses local representation structure to guide progressive token selection. Pruner applies the prior at both the token and neighborhood levels, using local dissimilarity as a multiplicative weight on token importance. We find that this prior is effective both on its own and in combination with existing token-importance signals. When combined with an existing spectral evolution score, Pruner retains 96.4% of the unpruned model's performance on average across eight multimodal benchmarks at an average visual-token budget of only 64 out of the original 576 tokens. This exceeds the previous state-of-the-art by 1.7 percentage points at the same token budget.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.