Many Caches, Same Reasoning: Rethinking Importance Selection for Reasoning-Time KV Compression
Abstract
KV-cache compression for long-form reasoning is usually framed as an importance-ranking problem: score past states, retain those judged most useful, and discard the rest. We ask whether successful retention policies actually agree on what those states are. Holding the full prefill cache fixed, we compare a signal-free page-level random retention baseline with R-KV, VaSE, and TriAttention across four reasoning models and three domains. Every compressed method therefore solves the same generated-state selection problem. Page-level random retention reaches 69.58% macro accuracy, defined as the unweighted mean across the 12 model-task cells, versus 70.45% for TriAttention, while exceeding R-KV and VaSE by 3.02 and 1.71 percentage points, respectively; its accuracy is within three percentage points of TriAttention in 11/12 model-task settings. More strikingly, independently successful random runs do not converge on the same cache. Across 2,000 multi-seed trajectories, their chance-adjusted overlap is (95% CI ) at comparable states, indistinguishable from independent random selection. Strong importance-based selectors likewise retain substantially different states, with adjusted pairwise overlaps of only 0.123–0.183; matched shared-core interventions reveal at most partial, asymmetric core effects rather than one common solution. These results suggest that reasoning-time KV compression is governed less by recovery of one privileged subset than by a broad space of distinct cache configurations that can preserve successful computation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.