To Predict or Not to Predict: Budget-Aware KV-Cache Compression
Abstract
Key-value (KV) caching avoids repeated computation during autoregressive decoding, but its memory cost grows with context length. Same-layer cross-head prediction can improve the low-rank compressibility of KV representations, but retaining anchor heads consumes cache capacity. Under a fixed KV-cache budget, prediction must therefore improve compressibility enough to justify storing the anchors. We propose TPNP (To Predict or Not to Predict), a KV-cache compression method that jointly determines whether to use cross-head prediction and how much of the prediction residual, termed the innovation, to retain. TPNP considers the native representation, prediction-assisted innovation coding, and prediction-free joint low-rank coding in a unified search space. When prediction is selected, TPNP retains same-layer anchors and jointly represents the target innovations in a low-rank subspace. TPNP ranks innovation basis directions using Fisher-Rao sensitivity and jointly selects anchor-target partitions and retained innovation ranks for each layer-wise K/V component to minimize the estimated Fisher-Rao compression cost under a global KV-cache budget. Our analyses reveal a key-value asymmetry: prediction generally improves compressibility in the value partitions selected by TPNP for cross-head prediction, while TPNP predominantly selects prediction-free joint low-rank coding for keys. We evaluate TPNP across four backbones and multiple calibrations. Across four KV saving ratios from 10% to 70%, comparisons with six methods show that TPNP ranks first or second in 14 of 16 backbone-ratio combinations for average perplexity and in 14 of 16 for average accuracy. On Qwen2.5-7B and Yi-1.5-6B, it also achieves the highest average LongBench score in six of eight combinations, ranking second in the remaining two.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.