acceptodds
Under review as a conference paper at ICLR 2027

Never Evict Always Taper: A Static Logarithmic Bit-Width Schedule for KV Cache Quantization

Abstract

Long-context inference relies on an ever-growing key-value (KV) cache, making precision allocation a key challenge for memory-efficient generation. We study how to distribute precision across cached tokens while preserving access to older information, without requiring online attention tracking. Under an exponential distortion surrogate and a prescribed precision floor, reverse water-filling yields an optimal allocation in which precision depends logarithmically on token importance; with power-law age weighting, this naturally produces a precision schedule that decreases with token age. A complementary perturbation analysis establishes a sufficient precision floor for preserving the highest-scoring attention key. Motivated by these results, we propose TaperKV, a token-preserving KV cache with a static, age-indexed integer-precision schedule that is compatible with a broad range of quantizers. Experiments on retrieval, language modeling, and mathematical reasoning show that TaperKV substantially reduces cache storage while maintaining strong task performance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.