LoopDraft: Target-Guided Multi-State Recurrence for Speculative Decoding
Abstract
Increasing a speculative drafter’s physical depth can improve proposal quality, but also ties computation to independently parameterized layers. A drafter need not match the target model’s full capability: it only needs to approximate short-horizon predictive distributions. This raises the question of whether the benefit of depth requires distinct layers or can instead be recovered through iterative refinement. We introduce LoopDraft, which reuses a shared computation block across an evolving latent state, recovering much of a deeper drafter’s benefit with about one quarter of its draft parameters. However, further recurrence quickly encounters diminishing returns, motivating better use of a fixed refinement budget. LoopDraft therefore maintains multiple persistent states and uses verified-prefix target features to guide state access and retention while preserving shared-core recurrent computation. On Qwen3-8B, R5S4 reaches 4.55× speedup with a mean acceptance length of 6.35, improving over a local five-layer DFlash baseline by 6.4% while reducing draft parameters by 73.9%. On Qwen3.5-4B, the same design improves acceptance length across DFlash, DSpark, and DFlash2.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.