Look Back to Think Less: Rethinking Visual Token Pruning Efficiency under Multimodal Reasoning Cost
Abstract
Visual token pruning aims to reduce the inference cost of vision-language models (VLMs). Existing methods typically retain a fixed subset of visual tokens during prefill, reducing computation through input compression. However, as VLMs evolve toward stronger reasoners, we observe an overlooked shift in computational burden from prefill to decoding. Discarded visual evidence and fragmented context often trigger repeated verification and multiple answer attempts that substantially inflate generation length. Motivated by this observation, we propose VISTA, a training-free framework that jointly controls visual evidence retention, dynamic access, and reasoning termination. VISTA constructs a compact and diverse visual reserve during prefill, dynamically revisits relevant evidence under a fixed visual budget based on attention trajectories, and uses answer consistency with lightweight periodic probes to terminate reasoning once a stable candidate emerges. We extensively evaluate VISTA on three VLMs across seven multimodal benchmarks covering document understanding, visual reasoning, and medical VQA. Under strict visual budgets, VISTA reduces average generation length by 36.3% while improving average performance by 25.4% over the strongest baseline, highlighting the potential benefits of jointly optimizing visual compression and subsequent decoding for end-to-end inference efficiency.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.