acceptodds
Under review as a conference paper at ICLR 2027

HEAT: Hierarchical Elastic Adaptive Token Sparsification via Depth Scaling

Abstract

Excessively long visual token sequences introduce substantial computational and memory overhead during the prefill stage of large vision-language models (LVLMs). Existing training-free visual token reduction methods alleviate this burden by pruning visual tokens at predefined LLM layers. However, fixed-depth pruning implicitly assumes that the same reduction depths are appropriate for all inputs, resulting in a mismatch between pruning depth and visual processing demand. To address this issue, we propose Hierarchical Elastic Adaptive Token Sparsification (HEAT), a training-free framework that scales pruning depth according to the evolving relevance state of each input. Specifically, HEAT introduces a cross-layer relevance monitor to quantify visual relevance concentration under the current token-retention budget, facilitating input-specific identification of appropriate pruning depths. Building on this, HEAT further develops a depth-adaptive threshold scheduler to adjust the pruning criterion according to both reduction ratio and model depth, supporting progressive visual token reduction at appropriate depths. Experiments across five LVLM backbones and a broad range of image, document, and video benchmarks demonstrate the effectiveness and generalizability of HEAT. With 66.7% of visual tokens removed, HEAT preserves 99.4% of full-token performance, reduces visual KV-cache memory by 66%, and accelerates LLM prefill by up to . Code is available at https://anonymous.4open.science/r/HEAT-C08E/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.