StacKV: Spatio-Temporally Aware Compression of KV Caches in Hybrid Language Models
Abstract
Hybrid language models, which interleave linear recurrent layers with full-attention layers, offer a promising balance between training efficiency and expressive power. However, during autoregressive generation, the key-value (KV) cache of the attention layers also poses a severe memory bottleneck. Existing KV cache eviction strategies, primarily designed for pure Transformers, are sub-optimal for hybrid architectures. They rely exclusively on signals from full-attention layers, without accounting for recurrent-state dynamics, and employ unconstrained global pruning, which disrupts the contiguous reasoning chains required for complex tasks. In this paper, we introduce StacKV, a training-free KV cache compression framework tailored for hybrid models, addressing both what to retain and where to retain it. First, we identify a complementary retention signal from recurrent-state dynamics: tokens with low-magnitude recurrent updates are often overlooked by attention-based KV selection, yet retaining them alongside high-attention tokens improves performance. Second, to prevent the compounding loss of local context during rolling eviction, we propose an age-aware non-uniform spatial blocking strategy. Extensive experiments on 5 tasks demonstrate that StacKV outperforms existing baselines, maintaining over 95% of full-cache performance on average for each model at a 25% retention budget, while achieving a 1.69 serving throughput speedup in Nano-vLLM.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.