acceptodds
Under review as a conference paper at ICLR 2027

PureKV: Plug-and-Play KV Cache Optimization with Spatial-Temporal Sparse Attention for Vision-Language Large Models

Abstract

Recent efforts compress and accelerate Vision-Language Large Models (VLLMs) by pruning KV caches of unimportant tokens. However, these methods typically rely on attention scores to estimate token importance, making them incompatible with efficient attention mechanisms such as sparse attention, or requiring additional computation to obtain attention scores. Moreover, existing methods overlook how sparse attention alters the information structure of the KV cache, thereby compromising the effectiveness of KV cache compression strategies. To address this issue, we propose PureKV, a plug-and-play KV cache compression framework that is fully compatible with efficient attention mechanisms. Our method utilizes lower layer attention scores to estimate the importance of high layers' KV cache, enabling active pruning without compromising accuracy. In addition, we have designed a Spatial-Temporal Sparse Attention (ST-SpAttn) module specifically tailored for video KV cache compression algorithms. This module combines spatial and temporal attention sparsity to improve the compression efficiency of KV cache optimization algorithms by purifying spatial noise and temporal redundancy in KV cache. Meanwhile, ST-SpAttn also accelerates the prefilling stage of VLLMs. Extensive experiments on VLLMs have shown that PureKV achieves 10.0 × KV cache compression and 3.35 × prefill acceleration, with negligible quality degradation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.