acceptodds
Under review as a conference paper at ICLR 2027

TraceKV: Signed Error Feedback for KV-Cache Quantization

Abstract

Low-bit key-value (KV) cache quantization reduces the storage requirements of large language models, but quantization errors can alter their output distributions and generation trajectories. Because errors across tokens can cancel or reinforce one another, reducing individual reconstruction errors alone does not directly control their combined effect on generation. We propose TraceKV, a method that uses self-generated traces to guide cache reconstruction through signed error feedback while preserving the base quantizer's encoding format, bit widths, and metadata footprint. TraceKV generates a short continuation with the bfloat16 (BF16) model, then replays it to compute gradients of multi-step Kullback-Leibler (KL) divergence with respect to cached values. Within each layer and attention head, it accumulates signed projections of quantization errors onto these gradients and selects between independent quantization and residual-feedback reconstruction for each token, encouraging positive and negative contributions to offset. Trace replay then selects among the completed candidate caches. Across four quantization representations, TraceKV reduces cumulative trajectory distortion. On PatternKV and OSCAR, gains persist for 64 tokens beyond the H64 selection window; paired ablations support retaining error signs. In two evaluations with frozen data and protocols, generation restarts from the compressed prompt. With KVarN as the base quantizer, TraceKV increases the mean length of the generated prefix that exactly matches the BF16 continuation by 14.4 and 22.8 tokens over independent KVarN, respectively, with prefix length capped at 1024 tokens.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.