acceptodds
Under review as a conference paper at ICLR 2027

Many Caches, Same Reasoning: Rethinking Importance Selection for Reasoning-Time KV Compression

Abstract

KV-cache compression for long-form reasoning is usually framed as an importance-ranking problem: score past states, retain those judged most useful, and discard the rest. We ask whether successful retention policies actually agree on what those states are. Holding the full prefill cache fixed, we compare a signal-free page-level random retention baseline with R-KV, VaSE, and TriAttention across four reasoning models and three domains. Every compressed method therefore solves the same generated-state selection problem. Page-level random retention reaches 69.58% macro accuracy, defined as the unweighted mean across the 12 model-task cells, versus 70.45% for TriAttention, while exceeding R-KV and VaSE by 3.02 and 1.71 percentage points, respectively; its accuracy is within three percentage points of TriAttention in 11/12 model-task settings. More strikingly, independently successful random runs do not converge on the same cache. Across 2,000 multi-seed trajectories, their chance-adjusted overlap is (95% CI ) at comparable states, indistinguishable from independent random selection. Strong importance-based selectors likewise retain substantially different states, with adjusted pairwise overlaps of only 0.123–0.183; matched shared-core interventions reveal at most partial, asymmetric core effects rather than one common solution. These results suggest that reasoning-time KV compression is governed less by recovery of one privileged subset than by a broad space of distinct cache configurations that can preserve successful computation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.