acceptodds
Under review as a conference paper at ICLR 2027

How Does KV-Cache Eviction Shape Output Divergence?

Abstract

KV-cache eviction reduces inference memory while altering subsequent generation. Across the interventions studied, lower cumulative token mismatch with full-cache continuations is associated mainly with later and less frequent first divergence; exposure-weighted post-divergence mismatch rates are also lower but remain high. Under a specified stepwise maximal coupling, we decompose the expected mismatch fraction into a first-mismatch term plus post-divergence exposure multiplied by its mismatch rate. Across an exploratory cohort and an independent 90%-retention cohort of two language models, exposure accounts for approximately 85–92% of six pooled mismatch-gap point estimates, including matched-byte selector comparisons at 50% and 90% retention. In the 90% cohort, prespecified comparisons within fixed early-divergence cohorts show higher total variation in late than early post-divergence windows for both models and selectors. We characterize the sharp range of cumulative mismatch compatible with an exact profile of finite divergence-aligned observations over unrestricted autoregressive kernel pairs and construct unbiased estimators via conditional Monte Carlo.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.