acceptodds
Under review as a conference paper at ICLR 2027

PRESTO: Prefix-Aligned Tree Drafting for Diffusion Speculative Decoding

Abstract

Diffusion LLMs (dLLMs) have emerged as effective draft models for speculative decoding (SD), generating a block of draft tokens in a single forward pass. However, existing diffusion drafters follow a single draft path, even though dLLMs produce multiple candidates at every position, inducing a large combinatorial space of paths; exploring only a tiny fraction of it fundamentally limits acceptance length. We introduce tree-based drafting for diffusion drafters to exploit this multi-candidate structure, but find that naive tree drafting is suboptimal: diffusion marginals are inherently prefix-blind, mismatching prefix-based AR verification and yielding unreliable path ranking. We propose PRESTO (PREfix-aligned Scoring and priority-based Tree search for diffusion speculative decOding), a principled framework built on two principles: (1) candidate ranking should align with the prefix-based nature of AR verification, and (2) tree construction should prioritize paths with high verification potential to maximize acceptance length. PRESTO exhibits strong generalizability across diffusion drafting paradigms, applying unchanged as a drop-in framework to both dedicated diffusion drafters and self-speculative dLLMs that draft and verify within a single forward pass. Across diverse benchmarks, PRESTO achieves average end-to-end speedups of 1.5× over the diffusion drafter DFlash and up to 2.14× over self-speculative dLLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.