On the Limits of KV Cache Merging
Abstract
Merging is an appealing alternative to evicting tokens from the key–value (KV) cache of a long-context language model: rather than discard an entry, fold it into one that is kept. Existing KV-cache merging methods broadly follow two strategies: directly folding an evicted entry into a retained key–value slot, or augmenting the merge with a correction for the attention mass represented by that slot. We examine both and show that their headroom is limited, while a larger opportunity lies in deciding what to evict. We first consider merges that leave keys untouched and only edit values. We prove that, for keys in general position, such a merge can be exact for every future query only when evicting the token would already have lost nothing, and derive how close two keys must be for an approximate merge to beat eviction. On four rotary models from three families, even with each evicted token paired with its best possible kept entry, at most a few percent meet this condition, including in text repeated verbatim: contextualisation and the rotary embedding keep the keys of repeated tokens apart. We then consider compensated merges, which add a per-slot attention-mass correction and escape the theorem, yet under a matched protocol no published operator we port shows a resolved LongBench gain over its own eviction baseline; those that rewrite retained keys lose, and in the clearest case the loss disappears once the same merges leave the keys in place. Our own construction keeps every key fixed, fits the value mixture and the mass correction jointly against the attention output, and keeps only the merges that lower it; its gain is +0.70 on the 400-sample development split, where it resolves, shrinks to +0.18, unresolved, on the 800-sample held-out split, and reverses to −0.74 on a second model, Qwen2.5-7B-Instruct. Merging is thus bounded by key geometry. The more effective strategy is to improve which tokens are evicted rather than how they are merged: since eviction error is the evicted attention mass times a value gap, weighting attention by value norm gives a selector indistinguishable from the strongest eviction methods on LongBench under a shortened-generation protocol, where margins of a few tenths of a point are below what the benchmark resolves, and better on needle retrieval.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.