Are We Evaluating VLM Token Pruning Right? Looking Beyond Accuracy and Token Count
Abstract
Visual token pruning aims to reduce inference cost while preserving the capabilities of the original VLM. Yet existing evaluations rely on aggregate task accuracy and retained token count as proxies for pruning quality and actual efficiency. We reveal two limitations: 1) Aggregate accuracy can obscure whether a pruning method truly identifies critical visual tokens—many examples remain solvable under arbitrary token removal and models can exploit textual shortcuts when visual evidence is degraded. 2) Retained token count is an unreliable proxy for efficiency, as identical token budgets can incur different end-to-end costs depending on where and how pruning is performed. Consequently, comparisons based on aggregate accuracy and retained token counts fail to recover the global performance–efficiency Pareto frontier across methods. We introduce PAA-Bench, a ***plug-and-play*** evaluation framework for existing benchmarks, to evaluate whether pruning succeeds through meaningful token selection and whether this success translates into real speedup. To determine when token choice actually matters, we propose *Random-Pass@* to measure how often an example remains correct under random pruning, enabling comprehensive evaluation across different levels of selection difficulty. To compare efficiency fairly, we introduce *Preservation-Constrained Speedup* (PCS) to report the maximum end-to-end speedup each method achieves under a shared preservation target. Experiments across *8* methods, *3* VLM backbones, and *9* benchmarks show that stronger token compression does not necessarily yield greater speedup, method differences widen on examples where token selection matters, and efficiency rankings change with the preservation target.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.