Specula: Lossless KV-Cache Recovery via Precision-Dual Drafting
Abstract
Disaggregated LLM serving separates prefill and decode across nodes, exposing generation to a critical vulnerability: decode-node failures destroy the KV cache and force serving systems to choose between fast-but-lossy compressed loading and slow-but-exact full loading. The dominant recovery strategy treats these as the only two options, implicitly assuming that exactness and speed are fundamentally opposed. In practice, this assumption misses a structural property of the setting. An INT4 checkpoint and its FP16 original are not two independent approximations. They encode the same generation history at two levels of precision, built from identical weights and prompt history. We call this precision duality, and it enables recovery that is both immediate and exact. We present Specula, a speculative recovery framework: the INT4 checkpoint serves as a draft for immediate generation while the exact FP16 state loads in the background, and rejection sampling verifies every committed token against the exact state as it arrives. Per-channel INT4 quantization preserves the variance structure governing argmax stability, yielding 97–99% acceptance across all tested models without tuning. Across five models from 7B to 32B parameters at 8K to 64K context, Specula delivers its first retractable token at least before the strongest lossless baseline commits, sustains continuous output through the recovery window where single-draft and lossless variants deliver one token or none, and shows no measurable quality degradation on HumanEval, GSM8K, and MMLU. By reframing recovery as a precision-duality problem rather than a speed–quality trade-off, Specula suggests a new perspective on fault-tolerant LLM serving.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.