The KV-Surgeon: Compensation-Aware Offline KV-Cache Compaction
Abstract
Offline compaction of the key-value cache (KV-cache) reduces the cost of repeatedly querying a fixed context by constructing a smaller cache once and reusing it subsequently. The central challenge of compaction is deciding which cache entries to delete and which to retain. Existing methods score each candidate for deletion on its own, with the rest of the cache being held fixed. However, softmax renormalization of the attention operation couples these decisions in a nonlinear manner: removing one entry changes the attention weights of every remaining one. In this work, we introduce KV-Surgeon, which uses calibration data to take interactions into account in a tractable way. We derive a quadratic surrogate that admits closed-form expressions for jointly making the optimal greedy removal decision as well as compensation within the remaining entries. Across benchmark datasets, such as QuALITY, LongHealth, and Long Code Arena, and at every evaluated compaction ratio, KV-Surgeon substantially advances the cost-accuracy Pareto frontier of dense and mixture-of-experts models from 4B to 30B parameters. We observe gains of up to four accuracy points on question-answering at compaction, exceeding the state of the art, attention matching (Zweiger et al., 2026), at comparable cost. Lastly, we introduce an even faster variant, which also exceeds the state-of-the-art while requiring only two fifths of its cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.