acceptodds
Under review as a conference paper at ICLR 2027

RFlash: Relaxed Speculative Decoding with Drift Constraint

Abstract

Speculative decoding losslessly accelerates memory-bound decoding of large language models (LLMs) by using a lightweight draft model to speculate future tokens of the large target model. Relaxed speculative decoding further accelerates decoding by accepting draft tokens that lossless verification would reject. Despite the increase in acceptance length, prior methods often fail to deliver any request-level speedup because they make the output longer. We trace this to a systematic drift toward the weak draft model, which accumulates over the generation and eventually collapses it. To address this, we propose RFlash, a training-free relaxation rule that caps the drift toward the draft, keeping long generations faithful to the target. On long-reasoning benchmarks, RFlash is the only relaxation rule we evaluate that is faster than lossless speculative decoding in every setting, by 6–10%, with output length and accuracy at the lossless level, delivering a practical, accuracy-preserving speedup.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.