WriteKV: Query-Agnostic KV Cache Eviction via Native Writes in Hybrid Attention Models
Abstract
During autoregressive inference, transformer-based large language models (LLMs) cache the key and value (KV) representations of previously processed tokens to avoid recomputing historical representations. In hybrid-attention models, linear-attention layers maintain a fixed-size recurrent state, whereas full-attention layers retain an ever-growing KV cache whose memory and per-step attention costs increase with the history length. This coexistence suggests that the update dynamics of linear-attention memory may provide a query-independent signal for compressing the full-attention KV cache. We observe that the native writes generated when linear-attention layers update their fixed-size states are conditioned on the current memory contents: writes aligned with existing memory directions are more likely to be redundant, whereas writes that introduce distinct directions are more likely to encode new memory content. Thus, the novelty of a write direction relative to the existing memory can serve as a retention signal for full-attention KV entries before future queries arrive. Based on this observation, we propose WriteKV, a query-agnostic KV eviction method that quantifies the novelty of write directions using blockwise Ridge statistics and physically compresses the full-attention KV cache through global layer-token selection. WriteKV requires neither additional context reconstruction nor a trained scoring model, and the resulting cache can be reused by multiple queries over the same history. Across three multi-query long-context benchmarks, WriteKV outperforms all evaluated KV-cache eviction baselines under a 20% KV budget. In system measurements with an 800K-token prefix at the same budget, it reduces decode-time attention latency to 26.8% of the full-KV attention latency, while increasing cache-construction time by only 4.20% relative to directly pre-filling the full KV cache.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.