One Objective, One Stage: Unifying Relevance, Coverage, and Redundancy in Visual Token Pruning
Abstract
Multimodal large language models (MLLMs) process large numbers of visual tokens alongside text, incurring substantial computation and inference latency. Existing training-free visual token pruning methods typically rely on single heuristics such as relevance, coverage, or redundancy. Although some approaches mitigate the limitations of single criteria through multi-stage or multi-module combinations, they still lack a principled, unified measure of the retained set value. Therefore, we propose **UniPrune**, a unified visual token pruning method that operates under a single objective in a single stage. UniPrune first utilizes cross-modal relevance operators to measure how visual tokens support text semantics. Subsequently, it uses Log-Sum-Exp (LSE) aggregation to combine the complementary support provided by different visual tokens, while its logarithmic form reduces the marginal benefit of redundant support. This objective jointly and explicitly captures relevance, semantic coverage, and redundancy suppression. We prove that this objective is monotone submodular, allowing the classical greedy algorithm to provably achieve an approximately optimal solution with low complexity. Across multiple benchmarks on three MLLM base models, UniPrune consistently achieves the highest average performance retention among all compared methods, including under the most aggressive 95% visual token pruning. UniPrune is independent of the vision encoder, requires minimal architectural changes, and remains compatible with acceleration methods such as FlashAttention.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.