acceptodds
Under review as a conference paper at ICLR 2027

SIDEKV: SIDE-INFORMATION-AWARE CONTEXT COMPRESSION FOR KV CACHE

Abstract

Long-context inference in large language models (LLMs) incurs substantial memory and computation through the key–value (KV) cache. Learned context compression replaces a long historical block with a compact representation, but existing approaches typically construct this representation from the discarded context alone, despite correlated context remaining available to the predictor. We introduce SideKV, a side-information-aware context-compression framework that conditions the rate of a learned summary on the retained context. A lightweight encoder compresses the discarded block into summary vectors, while a conditional rate model exploits the retained context as side information; the pretrained LLM remains frozen. Across 13 LLMs, with 320 of 1312 cache slots retained, conditioning on side information reduces the variational rate to – of that of an otherwise identical unconditional model and lowers excess predictive distortion on 11 of the 13 LLMs. The distortion advantage also persists across additional corpora, rate–distortion settings, and correlated time series. Among learned compressed representations, SideKV substantially improves over mean pooling while avoiding adaptation of the pretrained predictor, and achieves a mean excess distortion of bits/token. It requires only full-context prefill arithmetic and time to first token at the main setting. On LongBench, SideKV also achieves the highest mean score among the learned compressed representations evaluated. These results show that exploiting retained context as side information is a promising direction for lightweight KV-cache compression.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.