DISENKV: DISENTANGLING SHARED AND QUERYSPECIFIC EVIDENCE FOR KV CACHE RETENTION
Abstract
Existing observation-based key-value (KV) cache eviction methods mainly optimize importance scoring, evidence aggregation, or budget allocation while largely taking the observation evidence itself as given. We show that recent queries contain strong shared directional structure rather than independent retention signals. This structure can repeatedly reinforce the same historical tokens and obscure evidence expressed by fewer observations, inducing common-demand bias. We introduce DisenKV, a training-free method built on pre-aggregation query-side evidence disentanglement. Query-specific residuals supply the primary evictionrisk evidence, while raw queries retain shared demand for contextual coverage and late complementary rescue. Leave-one-out projection implements this principle by removing each query’s component along a shared direction estimated from the remaining observations. The raw and residual views are aggregated separately before a one-sided correction supplements residual risk. Across 16 LongBench tasks and four model families from 3.8B to 24B, DisenKV improves the strongest reproduced non-layer-aware baseline on every model at 20% KV retention and closes 32.5% of the remaining average gap to Full Cache. Under an identical global layer-wise allocation rule and total cache budget, Layer-DisenKV again improves all four models and closes about 44% of the corresponding gap. Controlled counterfactual interventions show that amplifying shared query structure systematically increases support-associated retention dominance under raw scoring, whereas leave-one-out residualization suppresses this response. These results support disentangling retention evidence before aggregation as a complementary approach to improving importance scoring and budget allocation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.