acceptodds
Under review as a conference paper at ICLR 2027

Borrowing Diffusion’s Gaze for VLM Visual Token Pruning

Abstract

Vision–language models (VLMs) encode images into many visual tokens, most of which are query-irrelevant yet computationally expensive. Existing methods perform one-shot or layer-wise pruning based on visual sensitivity, feature redundancy, or text–visual attention. However, these signals are largely derived from the target VLM itself and may be affected by structural biases such as attention sinks, limiting their reliability as indicators of query-relevant visual evidence. We introduce a Layer-0 pruning framework that draws on the strong text–image alignment of pretrained diffusion models to improve query-relevant visual evidence selection. Specifically, we distill diffusion text–image attention into DiffScorer, a lightweight scorer that provides an external signal for selecting query-relevant visual tokens before LLM prefill. To preserve this strong distilled prior while correcting its residual attention biases, counterfactual Swap-GRPO explores budget-preserving token exchanges around DiffScorer's deterministic selection and optimizes the swap trajectories with answer rewards from a frozen VLM. Only DiffScorer is updated in both stages. Extensive experiments across VLM architectures, pruning ratios, and visual understanding benchmarks demonstrate that DiffScorer consistently outperforms competitive one-shot, layer-wise, and learned pruning methods. Notably, for high-resolution inputs containing over 10K visual tokens, DiffScorer achieves a improvement in end-to-end generation throughput while preserving 99.5% of the unpruned accuracy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.