acceptodds
Under review as a conference paper at ICLR 2027

IDEAL-KV: Inference-Aligned Codebook Learning and Calibrated Stage Allocation for KV-Cache Compression

Abstract

Residual-codebook quantization reduces key–value (KV) cache storage, but preserving prediction quality at high compression ratios depends on both the learned codewords and the allocation of residual stages. Reconstruction error alone does not fully characterize the predictive effects of these decisions. We propose IDEAL-KV, which combines inference-aligned codebook learning with continuation-calibrated global allocation. With the language-model backbone frozen, IDEAL-KV learns structured Key and Value codebooks by aligning attention maps, attention outputs, and continuation distributions with the dense model. It then measures the continuation KL induced by each layer–component–depth choice in an otherwise dense cache and uses the resulting additive scores to allocate residual stages under a global KV-payload budget. The selected offline policy retains all historical tokens and supports packed-code attention without materializing the full historical KV cache during decoding. Across three LLM backbones, IDEAL-KV achieves near-BF16 performance on LongBench, RULER-32K, and GSM8K at approximately historical-KV payload compression. At approximately , LongBench and RULER-32K averages remain within 0.67 and 0.80 points of BF16, respectively. Code is available at https://anonymous.4open.science/status/IDEAL-KV-7562.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.