QCPruner: Query-Conditioned Population Coverage for Visual Token Pruning
Abstract
The high visual-token load in multimodal large language models (MLLMs) motivates aggressive training-free pruning, but under a prescribed final token budget, deciding which visual evidence survives is critical to preserving downstream performance. Existing methods rank tokens, diversify selected subsets, or combine relevance with coverage, while the direct conditioning of visual-to-visual population coverage by contextual per-visual query utility remains underexplored. We introduce QCPruner, a training-free method that uses contextual query utility to condition population-level visual-token coverage. Using keyword-matched query anchors, QCPruner fuses two cross-modal cues into utility and incorporates it into a visual-affinity-based coverage objective, where utility modulates representative suitability and, in the full formulation, also weights target importance. The resulting nonnegative facility-location objective is monotone and submodular, retains the standard 1-1/e greedy guarantee, and requires no model training or parameter updates. Across LLaVA-1.5, LLaVA-NeXT, LLaVA-Video, and Qwen2.5-VL, QCPruner consistently preserves strong downstream performance under matched token budgets, with the clearest gains appearing under aggressive compression. At 32 of 576 tokens on LLaVA-1.5-7B, it retains 96.1% of unpruned performance, versus 93.9% for the strongest evaluated baseline. At 256 of 1296 tokens on Qwen2.5-VL-7B, the corresponding values are 96.7% and 92.5%. Role-wise ablations show that representative-side query conditioning accounts for most of the gain over visual-only coverage, while target-side weighting provides a smaller, setting-dependent refinement.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.