SOAR: Visual Token Pruning via Cross-Modal Subspace Overlap Guidance in VLMs
Abstract
Abundant visual information strengthens vision-language model (VLM) perception, yet massive visual tokens raise inference costs. Existing visual token pruning methods have achieved significant progress by leveraging inter-token similarity or attention scores. However, such pruning guidance is largely restricted to local token-to-token similarity, ignoring the global geometric and structural relationships between the visual and textual subspaces. In this paper, we revisit VLM inference from a geometric perspective and present a novel guidance scheme for token compression. In particular, we identify a key observation: as LLM layers deepen, the subspaces spanned by the visual tokens and the textual tokens gradually align, and their overlap expands. To quantify this phenomenon, we propose Cross-Modal Subspace Overlap (CSO) to measure the extent to which visual tokens are spanned by text representations, revealing that more visual tokens in deeper layers can be linearly predicted by the text subspace. We accordingly propose Cross-Modal Residuals (CMR), which projects visual tokens onto the text subspace via Tikhonov-regularized least squares and exploits reconstruction residuals as an efficient geometric proxy to score tokens based on their unaligned features. Finally, based on CMR, we present SOAR (Subspace Overlap and Alignment Residuals), a training-free visual token pruning framework. Experiments on diverse VLM architectures verify the effectiveness of SOAR. For instance, on LLaVA-NeXT-7B, SOAR keeps only of visual tokens while preserving of the original average performance, achieving a prefill speedup, end-to-end speedup, and a KV-cache reduction.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.