TFPrune: Pruning Visual Redundancy while Preserving Token Functions for Efficient VLMs
Abstract
Large vision-language models (LVLMs) achieve strong multimodal reasoning by converting visual inputs into long token sequences processed by every subsequent LLM layer. High-resolution inputs further increase sequence length despite spatial redundancy, resulting in considerable inference costs. To reduce this overhead, numerous pruning methods have been proposed, many of which estimate redundancy from attention saliency or feature similarity. However, these scores do not assess whether the downstream effects of removed tokens can be reproduced by the retained set. To address this limitation, we propose TFPrune, a training-free, plug-and-play framework based on future functional substitutability. Specifically, TFPrune introduces an instruction-conditioned functional score that uses an exact softmax leave-one-out identity to quantify the counterfactual effect of removing each visual token on subsequent attention outputs without requiring an additional decoder-block forward pass. It selects a subset that minimizes the maximum functional discrepancy between each unselected token and its nearest retained representative, thereby preserving evidence that is difficult to substitute. Rather than simply discarding unselected tokens, TFPrune transfers their features to functionally matched representatives and uses accumulated transfer weights to correct post-compression attention. Experiments on LLaVA-1.5-7B, LLaVA-NeXT-7B, and Qwen2.5-VL-7B across nine benchmarks and multiple compression ratios demonstrate its effectiveness. Applied to LLaVA-1.5-7B, TFPrune retains only 11.1% of visual tokens while preserving 98.3% of full-model performance, achieving a favorable trade-off between compression and accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.