RESET: Residual-Energy Subspace Expansion via Tokens for Efficient VLM Inference
Abstract
Vision-Language Models (VLMs) achieve strong visual understanding but incur substantial memory and inference costs due to long visual token sequences. While numerous token compression strategies have been explored, we revisit the problem from a representative subset selection perspective, asking which visual tokens can collectively preserve diverse information with minimal redundancy. Based on this view, we propose RESET, a training-free visual token compression framework based on residual subspace selection. For images, RESET iteratively selects tokens with the highest residual energy, favoring previously unexplained feature directions and reducing spatial redundancy. For videos, RESET constructs a compact historical reference via SVD and conditions current-frame selection on the subspace orthogonal to historical information, thereby suppressing temporal redundancy. RESET is text-independent, plug-and-play, and compatible with FlashAttention, enabling seamless integration into existing image and video inference pipelines without modifying the underlying attention mechanism. Extensive experiments across multiple VLMs and benchmarks show that RESET consistently outperforms strong state-of-the-art token compression methods, validating residual-subspace selection as a simple and effective principle for representative visual token compression.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.