Stitch-and-Repair Speculative Decoding
Abstract
Speculative decoding accelerates language model inference by having a small draft model propose a block of tokens that a large target model verifies in parallel. A rejection discards the rest of the draft. Preparing continuations in advance helps, but only when they anticipate the target's sampled correction. We introduce Stitch-and-Repair, a training-free method that makes these misses both rarer and cheaper, recycling rejected computation instead of restarting from scratch. In addition to deterministic candidates, Stitch-and-Repair prepares a stochastic wildcard root which, unlike a deterministic candidate, can be coupled to the target's correction. After a rejection, it maximally couples speculative decoding's residual correction distribution with the wildcard's distribution, so the correction lands on prepared work more often while both marginals stay unchanged. On a match, the prepared continuation becomes the next proposal; on a miss, Stitch-and-Repair bridges the correction to a retained donor continuation through a single fresh draft token. All reused tokens are re-verified within the target's regular verification pass, exactly preserving the target distribution with no additional target forward passes. Across Llama-3.1-8B, DeepSeek-R1-Distill-Llama-8B, and Qwen3-8B targets on HumanEval, UltraFeedback, AlpacaEval, and GSM8K, Stitch-and-Repair improves end-to-end throughput over speculative decoding with continuations prepared in advance by 4.7–8.4% on pooled workloads and by up to 12.2% on individual datasets. It is also 20.2–37.3% faster than a tuned EAGLE-3 on the two targets with official EAGLE-3 checkpoints. Against a control that prepares identical work but samples the correction independently, coupling alone improves throughput by 2.2% on average and by up to 5.3%. Rejected speculation need not be wasted: coupling routes corrections to prepared work, and bridging reconnects the rest to retained work.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.