acceptodds
Under review as a conference paper at ICLR 2027

Retention Geometry: Where the Evidence Sits Decides How KV-Cache Compressors Rank

Abstract

KV-cache compression methods are compared by one accuracy number per benchmark. That number cannot identify how two methods rank once their position-resolved accuracies cross, and the deployed compressors make crossings common: at a matched budget, what a method keeps, and so whether the model answers, depends on where in the context the evidence sits. We formalise this with a method's evidence retention profile and a benchmark's evidence-position prior : a reported pairwise comparison has the sign of for the depth-resolved accuracy contrast , and window templates and accumulated attention have profiles that cross the random-eviction null's by construction. On a controlled grid (three models, 8K–32K, ten presses, six matched budgets, up to 21 depths paired within each item, two deployment protocols; 692,960 scored answers), the contrast between published methods changes sign with depth in 1–8 of 15 pairs in each harder-task cell under query-aware compression, certified by a bootstrap band corrected over depths and pairs; the count holds when the model's own greedy output is scored instead of teacher-forced tokens, though most reversals involve the first or last depth. A monotone link from predicts a held-out method's position-resolved accuracy at 0.47–0.77, against 0.41–0.58 for depth-averaged retention, an advantage carried by the window templates under query-aware compression and shared by most methods under query-agnostic compression. On natural documents scored by RULER's contains-match, 160 items certify 2 published pairs, only at the largest budget, and 7 under query-agnostic compression; RULER's own 996 QA rows, where depth is observed rather than paired, certify none, and re-weighting the grid by a task's measured prior does not forecast that task's ranking: the result is non-identifiability, not a forecast. An audit of 20,081 benchmark items puts the mean first-occurrence depth of 12 localisable tasks between 0.20 and 0.53, and finds one sub-task whose mean moves by 0.44 across RULER's releases. We recommend depth-resolved reporting and release a one-prefill retention probe. All code and every scored record are in the supplementary material and will be made public at https://github.com/xxx/xxx upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.