Beyond Attention and Similarity: Performance-Grounded Visual Token Pruning
Abstract
In this paper, we propose a performance-grounded visual token pruning framework, termed PerfPruner, which estimates token importance directly from the likelihood change of correct responses. By explicitly measuring each token's contribution to generating the target response, PerfPruner eliminates the objective mismatch inherent in prior proxy-based pruning methods, including attention-based and feature similarity-based approaches. To enable efficient inference, we first annotate a small training set with likelihood-based supervision and then train a lightweight predictor to estimate token importance scores at test time. We further introduce a score-constrained determinantal point process (SDPP) to reduce redundancy among retained tokens and improve pruning efficiency. By aligning token scoring with the pruning objective, our method can remove a broader range of irrelevant tokens, including redundant tokens, unique distractors, and out-of-distribution visual noise, thus providing a more reliable criterion for token selection across diverse image resolutions, pruning ratios, and tasks. Extensive experiments on 12 benchmarks spanning six tasks—including safety, hallucination, general understanding, reasoning VQA, text-rich/OCR, and chart/diagram understanding—demonstrate that PerfPruner consistently outperforms existing pruning methods across various pruning ratios and image resolutions, establishing new state-of-the-art performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.