CAPE: COMPRESSED ATTENTION WITH PRESERVED ENTROPY
Abstract
Reasoning models answer hard problems by generating long chain-of-thoughts (CoT), the KV-cache grows with the reasoning length and becomes the binding memory constraint on inference. KV-cache compression aims to maintain the same model performance while keeping only a bounded number of cached entries. Eviction-based methods have made this practical: scoring cached keys by the at- tention they receive from recent decode queries retains much of the uncompressed accuracy at a small fraction of the cache. We show that under a reduced cache bud- get, Softmax attention over the remaining keys is sharper than the distribution the model is calibrated for at its logical length, and the mismatch widens as the trace grows. A second issue is the loss of the objective: because keys are scored by the attention of recent queries, we found the model becomes less certain of the ques- tion as generation length grows under this eviction schema. We propose CAPE (Compressed Attention with Preserved Entropy), which restores the entropy of compressed attention to its uncompressed level, and reserves a bounded share of the budget for objective anchors selected by context reconstruction. Across three reasoning models on MATH-500 and AIME, CAPE recovers most of the accuracy lost under eviction with negligible additional cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.