Fork and Regenerate: MTP-Guided Early Forking and Suffix Regeneration for Block Diffusion Speculative Decoding
Abstract
Block diffusion accelerates speculative decoding by predicting multiple draft tokens in parallel. However, verification accepts only a matching prefix, so an early mismatch prevents the remaining draft tokens from being accepted. Moreover, later predictions cannot use the tokens selected at earlier masked positions in the same pass. To address these issues, we introduce MTP-guided early forking and suffix regeneration (FoRE). Since earlier predictions determine how much of a draft can be accepted, we investigate whether the target's pretrained multi-token prediction (MTP) module can improve candidate selection at the first drafted position. Our analysis shows that MTP provides higher target-token coverage than DFlash at this position, motivating us to retain multiple MTP candidates and supply each to the block diffusion drafter before generating the remaining tokens. To further improve acceptance at later positions, draft refinement selects alternative tokens and regenerates the remaining tokens from each modified prefix. Training with visible prefixes of varying lengths enables the same block diffusion drafter to perform both stages, without an added correction head or a separately trained second drafter. The resulting drafts are verified together in one batched target-model forward pass, so the output follows the target's greedy decoding. Across seven benchmarks on Qwen3.5-9B and Qwen3.5-4B, our method achieves 1.53–1.55× the average acceptance length of DFlash and 1.28–1.32× its average decoding speedup, reaching 3.82–3.93× over autoregressive decoding. When applied to DFlash 2, our method also improves acceptance length and speedup.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.