Speculating on the Past: Lossless Reuse of Stale KV Caches via Self-Speculative Verification
Abstract
Non-append edits in agentic LLM serving invalidate the cached suffix, forcing a choice between exact refresh, which delays generation, and approximate reuse, which can change the generated action. StaleSpec instead treats the position-aligned stale cache as a self-speculative proposal: the model decodes from the stale view while the exact post-edit suffix is recomputed in parallel, then commits only after fresh verification. The committed distribution is identical to refreshed-context decoding regardless of proposal quality; staleness affects only acceptance and latency. Across controlled 8K–32K edited turns, StaleSpec reduces committed latency by –, and under KV-packed prefill–decode serving it improves queue-inclusive throughput by – across three MoE deployments. A cost-aware gate selects profitable turns, and StaleSpec composes with lossless decode-side acceleration. Approximate stale reuse changes – of tokens and preserves only – of next actions, while StaleSpec remains exact.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.