acceptodds
Under review as a conference paper at ICLR 2027

LazyRecall: Decision-Risk-Aware Online KV Cache Management for Long-Form Generation

Abstract

Long-form generation and extended reasoning continuously expand the key–value (KV) cache, making GPU memory a bottleneck for scalable inference. Managing a fixed GPU KV budget requires three coupled decisions: what to evict, when to evict, and what to do if an evicted KV becomes relevant again. Existing methods often select KVs using attention statistics or local removal effects, which may not provide a globally comparable measure of their impact on the final model decision. Even with a precomputed eviction order, frequent rescoring incurs overhead, while bulk eviction removes states before their capacity is needed. Moreover, once evicted states are permanently discarded, subsequent rescoring cannot recover them if their relevance returns. We introduce LazyRecall, an online KV-cache management framework that jointly addresses selection, execution, and recovery. Decision-Risk uses a shared vector–Jacobian product (VJP) to approximate each KV’s effect on the final model decision, enabling model-wide selection on a common risk scale. On-Demand Removal decouples periodic planning from on-demand execution, retaining KVs until their capacity is required. Asynchronous Recall keeps evicted states recoverable on the CPU and retrieves relevant history for subsequent re-evaluation without blocking current decoding. Experiments on long-form writing and mathematical reasoning show that LazyRecall reduces the active GPU KV cache to as little as 15.5% of FullCache while maintaining comparable generation quality, with negligible additional per-request decoding latency. Anonymous code is available at https://anonymous.4open.science/r/LazyRecall/.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.