RESCUE: Gated Correction for KV Cache Eviction
Abstract
KV cache eviction is essential for efficient long-context LLM inference. Existing policies score KV entries from evidence already visible at the eviction point, and share a blind spot: at an aggressive budget each misses most of what the continuation will attend to, and policies reading the same observation window fail on overlapping entries rather than independent ones. We propose RESCUE, which treats future-aware eviction as correcting a deployed policy rather than replacing it and then verifying the correction against the dense model before it is applied. A policy-conditioned residual scorer is supervised only on the entries its base policy missed, so it never relearns what that policy already ranks correctly; a verification step then decodes a probe from the uncompressed cache immediately after prefill and adopts the corrected cache only if it keeps the output distribution closer to the dense model's. Verification is the part that carries the result: applied unconditionally the correction is harmful on almost a quarter of the evaluated cells. Across all 80 cells of 16 LongBench tasks × five base policies on Llama-3.1-8B, RESCUE improves the average score of every policy it augments, by +1.28 points (95% CI [+0.51, +2.16]), and cuts the cells it harms from 18 to 3; the average stays positive on Mistral-7B (+1.08) and Qwen3-8B (+0.43), with the gain concentrated on the weakest base policy and falling towards zero for the strongest as the budget widens. RESCUE targets evict-once-after-prefill inference and buys that safety with one dense decoding step.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.