Mind the Evidence: LLM-Aligned Visual Token Compression for VLMs
Abstract
Vision-language models increasingly rely on large numbers of visual tokens, making visual token compression essential for efficient inference. However, existing token pruning methods typically estimate token importance from the vision encoder or from attention/relevance patterns, without explicitly considering how visual information is transformed before reaching the language model. This creates a fundamental mismatch: a token that appears important in the visual representation may contribute little to the representation actually processed by the LLM. We introduce LAVIC, a training-free framework for compressing visual tokens from the perspective of the LLM-facing representation. LAVIC measures how distinguishable local visual variations remain in the LLM-facing representation, using this distinguishability as a signal of the visual evidence exposed to the downstream language model. Based on this evidence, LAVIC preserves informative tokens, groups redundant evidence around query-relevant anchors, and adaptively allocates the limited token budget according to the concentration of the evidence. Across multiple vision-language models and benchmarks, LAVIC consistently preserves model performance under aggressive token compression while substantially reducing visual tokens and inference cost. On LLaVA-v1.5, LAVIC retains 96.4% of the full-token performance using only 60 of 576 visual tokens, while reducing prefill latency from 58.70 ms to 28.95 ms. At 128 and 64 tokens, LAVIC also maintains strong performance and consistently outperforms competing compression methods. These results suggest that visual token compression can be better understood as preserving LLM-facing visual evidence rather than simply selecting visually salient tokens.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.