acceptodds
Under review as a conference paper at ICLR 2027

ERASE: Eliminating Redundant Visual Tokens via Adaptive Two-Stage Token Pruning

Abstract

Recent advances in Vision-Language Models (VLMs) enable large language models (LLMs) to process high-resolution images, significantly improving real-world multimodal understanding. However, this capability introduces a large number of vision tokens, incurring substantial computational overhead. To mitigate this issue, various vision token compression methods have been proposed. Existing methods often estimate different sources of visual redundancy using learned representations or fixed pruning schedules. We propose ERASE, an adaptive two-stage framework that separates image-level redundancy removal from instruction-dependent token pruning. Stage 1 derives image-dependent token retention from lightweight raw-image statistics, while Stage 2 progressively removes instruction-irrelevant tokens across decoder layers. Experiments demonstrate substantial token reduction while preserving accuracy: on Qwen2.5-VL-7B, ERASE retains 95.70% of the original model's accuracy at 25% token retention.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.