acceptodds
Under review as a conference paper at ICLR 2027

Ansight: Answer-Sensitive, Spatially Diverse Visual Token Pruning

Abstract

Modern vision-language models process hundreds to thousands of visual tokens for every query, making inference expensive in both computation and memory. Token pruning methods aim to reduce this cost, but the proxies they commonly use for token importance are not explicitly designed to capture the visual representations that actually matter for the model's answer. We show that the sensitivity of a VLM's answer to each visual representation can provide effective supervision for token selection. We identify Gradient Activation as an effective label-free signal derived from the frozen model's own generated answer, and distill it into Ansight, a lightweight question-conditioned selector that operates before the language model. However, relevance alone can spend a small token budget on a single spatial region or on redundant evidence. We therefore combine Ansight with regional spatial coverage and global feature diversity, retaining visual evidence that is compact but complementary. Extensive experiments across several backbones, benchmarks, and retention budgets show that Ansight achieves the highest average relative performance in 13 of 15 backbone–budget settings, reducing the performance loss of the strongest prior method by 26%. On LLaVA-NeXT-7B, it retains 94.9% of full-model performance with only 5% of visual tokens, while cutting computation by 8.6, reducing KV-cache memory by 94%, and accelerating end-to-end inference by 2.7. We further show that Ansight transfers to long-video question answering without video-specific retraining and extends to multi-turn dialog over a fixed image.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.