acceptodds
Under review as a conference paper at ICLR 2027

A Surrogate Distribution Perspective on Key-Value Cache Merging

Abstract

Merging entries in the key-value (KV) cache of large language models is a promising approach for managing GPU memory and throughput, and hence scaling to larger context windows. However, naively merging entries changes the scaled dot-product attention output at each layer, and if done aggressively seriously degrades accuracy on retrieval and reasoning tasks. In this paper we observe that the exact contribution of a merged cluster to attention is the cumulant generating function of the distribution of its constituent keys, together with the mean value under the exponential tilting of that distribution induced by the query. Compensating for a merge is therefore a question of which surrogate distribution to substitute for the cluster, and we show that a two-atom surrogate matching the cluster's first two key-norm moments yields closed-form corrections to both the attention score and the merged value, at a cost of two scalars and one vector per merged cache slot. We include in our analysis the Gaussian surrogate implicit in a second-order Taylor expansion of the attention score function. The resulting corrections are exact for two-token merges and for every query, and reduce exactly to standard attention for entries that have never been merged. Moreover, the statistics that our corrections depend on are invariant to the order in which merges are performed, so clusters can grow online as new tokens are processed. Experiments validate our theory by measuring fidelity of the induced corrections and sweeping performance on standard reasoning and retrieval tasks against merge threshold.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.