DEdit: Iterative Draft Editing for Speculative Decoding
Abstract
Speculative decoding accelerates autoregressive LLMs by having a lightweight drafter propose tokens that the target model verifies in parallel. Diffusion-based drafters further reduce drafting latency by proposing all tokens at once. However, these tokens are predicted independently, so a single early error causes prefix verification to discard the rest of the draft, even when it contains useful downstream predictions. We introduce DEdit, a diffusion-based drafter that can not only draft by conventional parallel unmasking but also iteratively edit its draft through token-to-token predictions. Through editing, later predictions can serve as bidirectional context for repairing earlier errors and extending the accepted prefix. To teach the model to repair errors while preserving correct predictions, we propose ProposalMix, a training scheme that mixes draft predictions with ground-truth tokens based on first-pass confidence during training. Across seven benchmarks, DEdit achieves the highest macro-average token acceptance among the evaluated drafters for both Qwen3-4B and Qwen3-8B under greedy and stochastic decoding, and its acceptance increases with more editing iterations and wider drafting windows. In our execution setting, speedup follows the same trends, and DEdit reaches the highest macro-average speedups of 5.72x and 5.97x over autoregressive decoding under greedy decoding.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.