Merge Then Drop, Adaptive KV Cache Compaction for Long-Context Reasoning
Abstract
As Large Language Models (LLMs) generate increasingly long reasoning traces, the Key-Value (KV) cache grows linearly and becomes a memory bottleneck. Learned KV Cache compactors typically assign the same budget to every layer and head, even though heads differ in how much context they need. We propose an online KV cache compactor that jointly learns the compressed representation and its per-head budget. The compactor produces an overcomplete set of merged candidate latents, from which differentiable Gumbel-Sigmoid gates decide which to retain under a global memory budget. Although gate allocation is fixed at inference, the compactor learns to route the relevant reasoning context into the surviving latents. Because backpropagating through the full sequential objective incurs an intractable memory cost for long contexts, we introduce an unbiased dual-sampling gradient estimator that retains only one compactor graph per trajectory. We evaluate on Qwen3-1.7B and 8B at compression ratios up to 16. On 8B, our method lies on the accuracy-memory and accuracy-latency Pareto frontier for math reasoning, improves AIME Avg@4 by up to 18.3 points over a uniform-budget compactor and matches a state-of-the-art learned eviction baseline at about 350 MiB of peak KV cache.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.