acceptodds
Under review as a conference paper at ICLR 2027

Speculating on the Past: Lossless Reuse of Stale KV Caches via Self-Speculative Verification

Abstract

Non-append edits in agentic LLM serving invalidate the cached suffix, forcing a choice between exact refresh, which delays generation, and approximate reuse, which can change the generated action. StaleSpec instead treats the position-aligned stale cache as a self-speculative proposal: the model decodes from the stale view while the exact post-edit suffix is recomputed in parallel, then commits only after fresh verification. The committed distribution is identical to refreshed-context decoding regardless of proposal quality; staleness affects only acceptance and latency. Across controlled 8K–32K edited turns, StaleSpec reduces committed latency by –, and under KV-packed prefill–decode serving it improves queue-inclusive throughput by – across three MoE deployments. A cost-aware gate selects profitable turns, and StaleSpec composes with lossless decode-side acceleration. Approximate stale reuse changes – of tokens and preserves only – of next actions, while StaleSpec remains exact.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.