acceptodds
Under review as a conference paper at ICLR 2027

The KV-Surgeon: Compensation-Aware Offline KV-Cache Compaction

Abstract

Offline compaction of the key-value cache (KV-cache) reduces the cost of repeatedly querying a fixed context by constructing a smaller cache once and reusing it subsequently. The central challenge of compaction is deciding which cache entries to delete and which to retain. Existing methods score each candidate for deletion on its own, with the rest of the cache being held fixed. However, softmax renormalization of the attention operation couples these decisions in a nonlinear manner: removing one entry changes the attention weights of every remaining one. In this work, we introduce KV-Surgeon, which uses calibration data to take interactions into account in a tractable way. We derive a quadratic surrogate that admits closed-form expressions for jointly making the optimal greedy removal decision as well as compensation within the remaining entries. Across benchmark datasets, such as QuALITY, LongHealth, and Long Code Arena, and at every evaluated compaction ratio, KV-Surgeon substantially advances the cost-accuracy Pareto frontier of dense and mixture-of-experts models from 4B to 30B parameters. We observe gains of up to four accuracy points on question-answering at compaction, exceeding the state of the art, attention matching (Zweiger et al., 2026), at comparable cost. Lastly, we introduce an even faster variant, which also exceeds the state-of-the-art while requiring only two fifths of its cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.