Still: Amortized KV Cache Compaction in a Single Forward Pass
Abstract
A bottleneck for longer-horizon deployments of LLMs is the memory stored in the KV cache. Iterative KV cache compaction can bound the KV cache growth, but existing approaches each have their tradeoffs. Selection and eviction methods are cheap but are bound to a less expressive token space, while synthesis methods utilize the continuous embedding space but often rely on costly per-context optimization. To balance both approaches, we introduce Still: a series of learned compressive maps at each layer of the KV cache, trained on a constructed MCQ dataset. We evaluate our KV compression method on Qwen and other open source models across long-document question answering, mathematical & scientific reasoning, and code generation. Still beats baseline methods at higher compression rates (8x to 200x), and on reasoning tasks can stay within 1.8 points of full-context performance on MATH-500 and GPQA-Diamond whilst using up to 3.8x less accumulated KV cache.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.