acceptodds
Under review as a conference paper at ICLR 2027

CAPE: COMPRESSED ATTENTION WITH PRESERVED ENTROPY

Abstract

Reasoning models answer hard problems by generating long chain-of-thoughts (CoT), the KV-cache grows with the reasoning length and becomes the binding memory constraint on inference. KV-cache compression aims to maintain the same model performance while keeping only a bounded number of cached entries. Eviction-based methods have made this practical: scoring cached keys by the at- tention they receive from recent decode queries retains much of the uncompressed accuracy at a small fraction of the cache. We show that under a reduced cache bud- get, Softmax attention over the remaining keys is sharper than the distribution the model is calibrated for at its logical length, and the mismatch widens as the trace grows. A second issue is the loss of the objective: because keys are scored by the attention of recent queries, we found the model becomes less certain of the ques- tion as generation length grows under this eviction schema. We propose CAPE (Compressed Attention with Preserved Entropy), which restores the entropy of compressed attention to its uncompressed level, and reserves a bounded share of the budget for objective anchors selected by context reconstruction. Across three reasoning models on MATH-500 and AIME, CAPE recovers most of the accuracy lost under eviction with negligible additional cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.