acceptodds
Under review as a conference paper at ICLR 2027

CORGI: Concept-Oriented Redundancy-Guided Image Token Compression for Vision-Language Models

Abstract

Visual token compression aims to reduce the computational cost of Vision-Language Models (VLMs) by removing redundant or non-informative visual tokens before or during multimodal processing. Existing visual token pruning methods suffer from two major limitations. First, their pruning decisions are semantically opaque: practitioners cannot determine what information the kept or discarded tokens represent or control generation by selecting which visual tokens should be preserved for a particular task or application. Second, many methods require layer-specific design choices and hyperparameter tuning, such as selecting the transformer layers at which pruning should be performed. We introduce the first interpretable and controllable visual token pruning framework for VLMs. Our method identifies the semantic meaning of visual tokens in natural language by leveraging the model’s shared vision-language representation space. This allows an easy, human-friendly way of selecting visual tokens for pruning according to their textual meaning. The resulting token selection is performed before the visual tokens are provided to the VLM, eliminating the need for additional layer-wise token dropping or pruning schedules. Experiments on image and video benchmarks across both natural and medical domains demonstrate that our method often outperforms the state-of-the-art while offering greater interpretability, explicit control, and a simpler pruning pipeline.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.