acceptodds
Under review as a conference paper at ICLR 2027

SWIFT: Second-order Weighted Importance For Tokens

Abstract

KV-cache eviction reduces the memory and bandwidth cost of long-context inference, but its effectiveness depends on identifying cached entries whose physical removal minimally changes model behavior. Existing methods typically address the physical removal perturbation, downstream sensitivity, and global budget allocation separately, making eviction scores difficult to compare across layers and KV heads. We introduce SWIFT, a task-free second-order KV-cache eviction method grounded in a self-fidelity objective. Using only the cached context, full-coverage observations define fidelity against detached outputs of the uncompressed model, while exact softmax-renormalized removal perturbations and matrix-free scalar probes estimate second-order output damage without materializing the Hessian. These globally comparable scores enable SWIFT to allocate a shared physical KV budget across layers, KV heads, and tokens. Experiments cover controlled interventions and comparisons across three 7–8B model families and four long-context benchmarks. SWIFT achieves Pearson and Spearman correlations of and with measured eviction damage and ranks first or joint-first in 22 of 24 primary model–benchmark–retention comparisons. At 20% retention, it provides physical KV compression, delivers a 9.7–14.1% steady-state decoding speedup, and realizes approximately 80% of the empirically attainable historical-KV latency reduction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.