Looking Beyond Tokens: Context-Aware Visual Token Compression for MLLMs
Abstract
Multimodal large language models (MLLMs) suffer from heavy computational and memory overheads due to the excessive number of visual patch tokens. Existing visual token compression approaches typically rely on raw token importance scores or raw feature similarities, leading to redundant budget allocation or invalid merging across distinct objects. In this paper, we argue that reliable visual token compression should be guided by multimodal context, where token relations are jointly characterized by spatial structure and text-conditioned semantics rather than isolated visual features. Motivated by this insight, we propose Context-aware Visual Token Compression (CVTC), a novel training-free compression paradigm for MLLMs. CVTC employs a cascaded context coarsening mechanism to construct a multimodal context graph, progressively integrating spatial topology and text-conditioned relevance to guide context-aware token merging. Based on the propagated multimodal context, CVTC performs greedy contextual merging to identify tokens sharing coherent contextual relations, while aggregating their original visual features to construct the compressed representation. Extensively evaluated on various benchmarks across different backbones, CVTC achieves a favorable balance among performance, efficiency, and robustness. The codes are available in supplementary materials will be publicly available.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.