ReDraft: Your Verifier Is Secretly a Depth Supervisor for Recurrent Speculative Decoding
Abstract
Speculative decoding accelerates LLM inference by having a cheap drafter propose token blocks that the target model verifies in parallel. Yet drafter depth is typically fixed by ablation and applied uniformly across blocks, despite substantial variation in draft difficulty: a per-block depth oracle gains accepted tokens over the best fixed depth, while the cheapest depth suffices for 40% of blocks. Standard halting heuristics fail to exploit this variation: convergence- and confidence-based exits collapse to maximum depth, reducing throughput by 20-27%, and their tuned thresholds transfer poorly across tasks, model scales, and training data. We instead exploit a supervision signal unique to speculative decoding: the verifier reveals each draft's accepted length without additional model computation. We introduce ReDraft, a learned policy that selects drafter depth for each block. ReDraft uses a -layer weight-shared core iterated to variable depth and a lightweight head that predicts each block's acceptance curve from verifier-derived labels. A renewal-reward stopping rule converts these predictions into depth choices using measured hardware costs, while leaving verification—and thus losslessness—unchanged. ReDraft plans once per block with only 1.8% overhead, and a regret bound shows that throughput loss depends on errors in the relative values of candidate depths. Across five benchmarks and three targets spanning two model families, a single recipe yields tuning-free policies within 4-6% of the per-task tuned optimum, with no statistically significant difference on GSM8K, while allocating computation similarly to the oracle. The recurrent drafter achieves a lossless speedup on GSM8K, matching the released -layer DFlash drafter with 40% of its parameters and half its training data.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.