acceptodds
Under review as a conference paper at ICLR 2027

ProSpect: Efficient and Consistent Self-Speculative Decoding

Abstract

In single-user and on-device settings, large language model (LLM) decoding typically runs at batch size one and is memory-bandwidth bound. Every forward pass streams the full model weights to emit a single token, leaving the GPU's arithmetic idle. Self-speculative decoding turns that idle arithmetic into tokens without adding any weights, by drafting the next block of tokens in parallel and verifying the previous block autoregressively in a single pass. There is no auxiliary drafter to train, keep aligned with the target, or hold in memory beside it, and no ceiling on draft quality set by a smaller model's capacity. We address two shortcomings of self-speculation. First, parallel drafts are sampled independently and can be individually plausible yet jointly incoherent. By front-loading the randomness of sampling, we generate consistent drafts without sacrificing parallelism. Second, a decode step cannot know how many of its drafts will be accepted, so a draft branch has to be appended for every possible acceptance count, which makes the query size quadratic in the draft length. By drafting only for the most likely acceptance counts, we reduce the query size to linear. The resulting decoder, ProSpect decodes faster than autoregressive Qwen3-8B on SpeedBench, outperforming speculative decoding with dedicated drafters DSpark and DFlash by % and %, and faster than autoregressive Gemma-4-31B, ahead of DFlash by %.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.