RecastPrune: Rethinking Token Selection and Information Retention in Visual Token Pruning
Abstract
Training-free visual token pruning reduces the number of visual tokens processed by multimodal large language models without updating model parameters. However, existing methods often rely on image-only importance signals that may be misaligned with the user query, greedily select tokens without revisiting earlier decisions, and permanently discard the information carried by unselected tokens. We propose RecastPrune, which aligns token importance with the user query, revises the initial token selection, and selectively fuses information from unselected tokens. RecastPrune first applies query-aware relevance calibration by combining CLS-to-patch attention with patch–query similarity from the frozen CLIP encoders. It then performs coverage-aware initialization followed by restricted local exchange. Finally, saliency-aware selective absorption fuses relevant, compatible information from unselected tokens into retained representatives. Across LLaVA-1.5 and LLaVA-NeXT, RecastPrune achieves the highest average accuracy relative to the full-token upper bound at every evaluated token budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.