acceptodds
Under review as a conference paper at ICLR 2027

Quantize the Keys, Audit the Values: Retained-State Certificates for KV-Cache Compression

Abstract

KV-cache compression aims to support long-context inference within limited memory while preserving output quality. How that memory is distributed matters because compression in different cache regions can affect the current attention output differently. Aggregate task scores assess overall quality but do not by themselves provide query-specific feedback for allocating memory within a compressed cache. We provide this feedback with a deterministic certificate that bounds attention error for each query, even after the original entries are discarded. It also gives the exact reduction in the bound that each increase in value precision would produce, which enables adaptive allocation while the cache stays compressed. This allocation lowers attention error by 30–67% under the same memory budget and reaches a calibrated value-error target with roughly two-thirds as much refinement memory as uniform allocation. In packed serving of two 7–8B models at 32K, adaptive allocation brings the next-token distribution 2.6–3.4 times closer to dense decoding in KL divergence than uniform allocation with the same memory. The design is also practical to serve. A fused kernel evaluates the certificate alongside attention with little additional arithmetic and reads the selected state directly from a packed cache. Its base representation uses about dense KV storage and serves a 70B model at 128K on one GPU.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.