acceptodds
Under review as a conference paper at ICLR 2027

KV-Kaizen: Learning Context-Adaptive Cache Compression Choices

Abstract

As the context size of text processed with an LLM grows, the size of KV caches can outstrip the memory allocated for the original model weights. This impacts LLM throughput negatively, since decoding is memory-bound and decode cost grows with cache size. Recent work has proposed to alleviate this bottleneck by discarding some of these tokens, keeping the most relevant bits and dropping the rest. *Eviction* introduces a tension, since a one-off decision to discard content may prove detrimental later. Instead, we focus on alternative choices that can lead to cache compression **without** evicting tokens. We achieve this by learning a selector that is able to produce, based on context, a per-layer cache configuration towards an overall compression budget. The selector operates along three axes: sharing one cache across layers (*depth*), caching at fewer bits (*precision*), or truncating the low-rank latent cache representations (*rank*). We call the resulting method KV-Kaizen, for the many small per-layer choices it compounds. We observe that these interventions taken independently and uniformly over all layers limit achievable compression because they degrade accuracy. Crucially, we observe that composing them *locally* and *adaptively* to the context can instead preserve accuracy while achieving large memory savings. At train time, a lightweight selector module is given access to the embeddings of the context sequence, and trained to pick *at each layer* a combination of these three axes, with a language modeling loss that is regularized by a budget constraint satisfaction loss. At inference, the selector runs once, before pre-fill. In evaluations on instruction following and reasoning tasks, our selectors reach the Pareto frontier of accuracy against cache size, against learning-free and *post-hoc* baselines. On long-context tasks, KV-Kaizen improves on eviction and can be composed with it, reaching a smaller decode-time cache on a B model while preserving accuracy. A cache size reduction incurs no accuracy degradation from B parameters up, and a compressed model is more accurate than a smaller uncompressed one with the same cache size. Together, these findings support pre-training large models and compressing them only afterwards.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.