QD-Prune: Query-Driven Vision Token Pruning for Efficient Video-Language Models
Abstract
Video-language models encode videos into hundreds of vision tokens, which dominate the input sequence and drive up inference cost. Training-free pruning methods such as FastV remove tokens with low attention scores, but we find that their pruning signal is query-agnostic: it is taken from the last text token and stays nearly identical across different questions about the same video (cosine similarity 0.94), so query-critical tokens are removed and causal and temporal reasoning degrade. We propose QD-Prune, a training-free method that scores vision tokens by query-to-vision attention (M1) together with temporal diversity (M2), and selects the pruning layer by attention entropy minimization (M3). It requires no trainable parameters and reuses attention weights already computed during the standard prefill pass. At a 20% keep ratio on NExT-QA with Qwen2.5-VL-7B, QD-Prune reaches 64.3% overall accuracy, outperforming FastV by +5.3% (McNemar p=0.030), with the largest margin on causal reasoning (+7.0%). On MVBench it reaches 43.8%, +7.7% over FastV (p<0.001), matching the full-token baseline (44.1%, p=0.817).
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.