acceptodds
Under review as a conference paper at ICLR 2027

GroundPrune: Representation-Preserving Visual Token Pruning for Grounding in VLMs

Abstract

Visual token pruning reduces the inference cost of vision-language models (VLMs), and existing methods are evaluated by whether downstream answer accuracy is preserved. We argue this criterion is insufficient: a pruned model can still answer correctly while its internal representation has already diverged from that of the original model. To measure this directly, we compare pruned and original models at two levels: the KL divergence between their next-token distributions on general VQA, and the IoU between the boxes that a linear probe, trained on the original model, regresses from the last prompt token's hidden state of the pruned and of the original model. Across FastV, SparseVLM, VisionZip, and Nüwa, methods that retain most VQA accuracy still incur substantial distribution shift and degraded box decodability, an erosion masked by VQA accuracy but amplified on referring expression comprehension (REC). We trace it to two causes: saliency mode collapse at the vision stage, where retained tokens cluster around a few dominant hotspots, and the choice of routing point in the decoder, where only a few attention heads can guide pruning without disturbing the original model. We propose GroundPrune, a training-free cascaded framework that addresses each in turn: Visual Evidence Preservation (VEP) selects feature-diverse tokens from a saliency-guided candidate pool, and Behavior-Preserving Decoder Pruning (BPDP) prunes inside the decoder using the attention of a single routing head, located offline as the one whose pruning keeps the model closest to the original, with no supervision at inference. Across eight REC subsets and eight VQA benchmarks, GroundPrune attains the lowest KL divergence and probe error among compared methods; at 25% retention on Qwen2.5-VL-7B it reaches 89.7% relative REC accuracy, 5.1 points above the strongest baseline, while keeping 96.1% relative VQA accuracy, with consistent gains on LLaVA-1.5-7B and LLaVA-NeXT-7B.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.