PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models
Abstract
Recent efforts compress and accelerate Vision-Language Large Models (VLLMs) by pruning KV caches of unimportant tokens. However, these methods typically rely on attention scores to estimate token importance, making them incompatible with efficient attention mechanisms such as sparse attention, or requiring additional computation to obtain attention scores. Moreover, existing methods overlook how sparse attention alters the information structure of the KV cache, thereby compromising the effectiveness of KV cache compression strategies. To address this issue, we propose PureKV, a plug-and-play KV cache compression framework that is fully compatible with efficient attention mechanisms. Our method utilizes lower layer attention scores to estimate the importance of high layers' KV cache, enabling active pruning without compromising accuracy. In addition, we have designed a Spatial-Temporal Sparse Attention (ST-SpAttn) module specifically tailored for video KV cache compression algorithms. This module combines spatial and temporal attention sparsity to improve the compression efficiency of KV cache optimization algorithms by purifying spatial noise and temporal redundancy in KV cache. Meanwhile, ST-SpAttn also accelerates the prefilling stage of VLLMs. Extensive experiments on VLLMs have shown that PureKV achieves 10.0 × KV cache compression and 3.35 × prefill acceleration, with negligible quality degradation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.