acceptodds
Under review as a conference paper at ICLR 2027

Pruning by Consequence: Intervention-Guided Visual Token Selection for Vision-Language Models

Abstract

Training-free visual token selection reduces the inference cost of vision-language models by retaining a small subset of image tokens. Existing selectors rank tokens with attention, similarity, saliency, or diversity signals—inexpensive proxies for the quantity that actually matters: whether the compressed model still behaves like the uncompressed one. We measure that quantity directly. Our method, Cluster Intervention Gain (CIG), scores a candidate subset by how faithfully it preserves the full model's answer distribution, maximizing the negative KL divergence between the full-token and selected-token answer distributions through real subset prefills and a cluster-level greedy search; no labels or training are required. Under a controlled exact-576 protocol on InternVL3-8B, Qwen2.5-VL-7B, InternVL3.5-4B, and Qwen3-VL-4B, CIG is the strongest architecture-portable selector in all 24 backbone-budget settings, and paired intervals resolve the advantage for every backbone at k ≤ 64. Its mean margin over the best heuristic grows from +0.99 percentage points at k = 192 to +11.57 percentage points at k = 8, where it still retains at least 93.3% of full-token accuracy with 72× fewer visual tokens. A batched variant stays within 0.65 percentage points of the sequential selector while cutting selection time by 1.6–4.1×, and at serving-scale batching the shorter KV cache lowers end-to-end latency by up to 17.9% on a memory-pressured 8B workload, selection cost included. Ablations attribute most of the gain to the intervention objective itself rather than to the clustering heuristics that make the search tractable.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.