acceptodds
Under review as a conference paper at ICLR 2027

Look Back to Think Less: Rethinking Visual Token Pruning Efficiency under Multimodal Reasoning Cost

Abstract

Visual token pruning aims to reduce the inference cost of vision-language models (VLMs). Existing methods typically retain a fixed subset of visual tokens during prefill, reducing computation through input compression. However, as VLMs evolve toward stronger reasoners, we observe an overlooked shift in computational burden from prefill to decoding. Discarded visual evidence and fragmented context often trigger repeated verification and multiple answer attempts that substantially inflate generation length. Motivated by this observation, we propose VISTA, a training-free framework that jointly controls visual evidence retention, dynamic access, and reasoning termination. VISTA constructs a compact and diverse visual reserve during prefill, dynamically revisits relevant evidence under a fixed visual budget based on attention trajectories, and uses answer consistency with lightweight periodic probes to terminate reasoning once a stable candidate emerges. We extensively evaluate VISTA on three VLMs across seven multimodal benchmarks covering document understanding, visual reasoning, and medical VQA. Under strict visual budgets, VISTA reduces average generation length by 36.3% while improving average performance by 25.4% over the strongest baseline, highlighting the potential benefits of jointly optimizing visual compression and subsequent decoding for end-to-end inference efficiency.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.