acceptodds
Under review as a conference paper at ICLR 2027

Shallow Pool, Deep Prune: Accelerating Vision-Language Models via Hybrid Token Reduction

Abstract

Vision Language Models (VLMs) process hundreds or even thousands of visual tokens to perform complex visual understanding and reasoning tasks. However, such long input sequences result in substantial computational overhead during inference. While a range of token reduction approaches have been proposed to accelerate VLMs, they generally encounter early-stage context loss, inadequate salient token retention and suboptimal pruning allocation, leading to significant performance degradation. To address these issues, we propose a novel hybrid token reduction method, termed HTR-VLM, motivated by the depth-dependent variation in compression objectives across VLM layers. Specifically, we introduce the FFN Token Pooling with Residual Approximation (FTP) module, which leverages residual connections to approximate full image-feature updates with pooled tokens, reducing shallow FFN computation while preserving full context. Building on this, Salient Token Retention with Multi-Criteria Scoring (STR) selects and preserve important tokens for full computation by integrating multiple cues. Finally, a Depth-Aware Token Pruning (DATP) mechanism is adopted to enable more precise token reduction in deeper layers. Extensive experiments demonstrate that HTR-VLM substantially reduces computational overhead and maintains higher model performance compared to state-of-the-art token reduction methods. We will release our code upon paper acceptance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.