acceptodds
Under review as a conference paper at ICLR 2027

Same Memory, Better Answers: Choosing KV Caches by Their Outputs

Abstract

KV-cache compression research asks which tokens to keep, and judges each compressor by the one cache it builds. This misses a larger lever. At a fixed memory budget, different compressors succeed on different questions: retrospectively choosing the best of their caches for each question matches or beats the uncompressed model while keeping one-tenth of the document. Common eviction scores—attention mass, reconstruction error, and hidden-state distance—measure what a cache preserves inside the model, not its next-token distribution. On question answering over research papers, retained attention mass ranks candidate caches worse than a coin flip. We select caches by their outputs instead: each candidate is scored by how far it moves the next-token distribution from the full cache's. The geometry of the unembedding explains what the output score measures: it weights internal changes by their effect on next-token probabilities, information that internal distances discard. With the same candidate caches and scoring passes, output-based selection picks better caches than internal-state features on question answering. It also beats fixed policies tuned on training data and transfers without retraining from 7B to 70B models and to 128K-token documents. Internal features lead on the synthetic needle-retrieval task, where the answer is a verbatim copy.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.