acceptodds
Under review as a conference paper at ICLR 2027

BellowsKV: One Latent Cache for Many Layers, Expanded on Demand

Abstract

Large language models (LLMs) are now widely utilized as the backbone of agents that reason over many turns, call tools, and pursue long-horizon tasks. To serve such a context without recomputing the past, a model stores the attention keys and values of all past tokens in a key-value (KV) cache. Each layer maintains its own cache and re-reads it at every generation step, so KV-cache memory can become the main bottleneck for GPU serving. Most compression methods operate within a single layer, by quantizing the cache entries, reducing the channel dimension, or evicting less important tokens from the cache. All three lose substantial accuracy at high compression ratios. A more recent line of work shares one compressed cache across adjacent layers to go beyond this limit. However, existing cross-layer methods also degrade sharply at high compression ratios. We find that the KV cache representations of adjacent layers are quite different, whereas the input hidden states are highly similar. We also find that, under the same memory budget, reducing the dimension of this shared latent loses far more accuracy than quantizing it. Based on these findings, we propose BellowsKV, a training-free method that maps the hidden states of n adjacent layers to one shared latent per token. At inference time, each layer reads its own keys and values from this latent through a layer-specific projection. We keep the latent at full dimension and compress it in the two other ways instead: BellowsKV-Quant stores the latent at low precision, which is nearly lossless because the latent has no dominant outlier channels; BellowsKV-Evict makes one eviction decision per block, which the layers can share because they rank token importance very similarly. Across Llama-3.1-8B, Mistral-7B and Qwen2.5-7B, BellowsKV-Quant retains about 99% of full-cache accuracy at 3.6 compression and 97–98% at 13–14, where the strongest existing cross-layer method retains at most 84% at 16.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.