ViSparse: Retrieval-Based Visual Attention Sparsification for Multimodal Reasoning
Abstract
Extended reasoning enables multimodal large language models to solve complex visual tasks, but caching and repeatedly attending to thousands of visual tokens during long reasoning traces incur substantial memory and computational overhead. Token compression offers a promising way to reduce this overhead. However, existing methods can be limited by restricted selection flexibility and insufficient adaptation to visual attention that varies across reasoning steps. We introduce ViSparse, a retrieval-based framework for visual-attention sparsification. ViSparse combines head-specific hierarchical cluster retrieval with coverage-adaptive refresh to adapt to attention patterns during reasoning, selectively loading relevant visual token states from CPU to GPU. It further exploits attention similarity across adjacent layers to share retrieval indices, reducing redundant retrieval while allowing each layer to compute attention using its own visual token states. Across three multimodal reasoning models and eight benchmarks, ViSparse achieves the highest average accuracy among the evaluated compressed methods for each model, while maintaining accuracy comparable to full-cache attention. Efficiency experiments show up to 2.262× decoding speedup and up to 70.1% lower peak GPU key–value cache storage. These results show that ViSparse improves multimodal reasoning efficiency while retaining on-demand access to all visual token states.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.