acceptodds
Under review as a conference paper at ICLR 2027

KVRecode: Lossless KV Cache Compression with Local Coding and Buffer Reuse

Abstract

Long-context generation places increasing memory demands on the key and value (KV) cache. However, compressing this cache losslessly requires repeated reconstruction before attention, making inference cost depend on both the stored representation and its runtime. To address this challenge, we introduce KVRecode, a lossless BF16 cache representation and runtime that combines complete local exponent palettes with reusable reconstruction buffers. KVRecode encodes exponents with fixed-width indices into page-local tables while preserving sign and mantissa bits, retaining pages in raw form when their exponent sets exceed the palette capacity. This construction enables exact recovery without calibration or per-value exception handling. As the cache grows, the runtime seals completed token segments and reconstructs successive layers into shared buffers, limiting reconstruction memory while preserving the existing attention backend. Across two 8B models and nine natural inputs at 4K, 8K, and 16K contexts, KVRecode reduces request latency for 512-token outputs by 10.45% relative to SplitZip under matched resident-cache execution. It also uses 15.19% less resident cache memory than Raw, including tails and reconstruction workspaces. KVRecode preserves cache contents and generated outputs exactly. These results demonstrate the benefits of coupling local coding with buffer reuse for efficient lossless cache access during generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.