acceptodds
Under review as a conference paper at ICLR 2027

Fail Fast, Win Big: Rethinking the Drafting Strategy in Speculative Decoding via Diffusion LLMs

Abstract

Speculative decoding efforts have largely optimized draft quality for enhanced performance, treating policy as a secondary focus. While effective, this approach can be costly and brittle, as speedups depend on continually improving drafter quality. We argue that drafting policy should be a first-class driver of performance. Our key insight is that intrinsic properties of emerging diffusion LLMs (dLLMs) make this shift both possible and effective. Their accuracy improves sublinearly with compute, yielding diminishing returns to quality, and unlike with autoregressive models, drafting cost is nearly independent of length, enabling long drafts at low marginal cost. Together, these properties favor policies that hedge against draft quality – dynamically allocating compute across the sequence based on heterogeneous decoding difficulties across the sequence – rather than attempting to maximize it. We present FailFast, a speculative decoding framework that operationalizes this perspective. FailFast allocates minimal compute to hard regions ("fail fast") while aggressively extending drafts in easy regions ("win big"), often accepting tens of tokens per step. Without any fine-tuning, FailFast provides lossless acceleration of autoregressive LLMs, achieving up to 4.9× speedup over standard decoding, 1.7× over a naive dLLM drafter baseline, and 1.7× over EAGLE-3 across diverse models and workloads.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.