Self-Speculation for Faster Reasoning Models
Abstract
Reasoning language models generate long chains of thought (CoT) before responding, and for long-form outputs such as code, generating the final response accounts for a substantial share of end-to-end latency. Speculative decoding can accelerate this phase, but existing drafters either copy text the model has already written, as in prompt lookup, or use a cheaper approximation of the model, like a small draft model, to guess a few tokens ahead. Reasoning models expose an additional source of speculation: complete responses produced by the model from a partial CoT, which increasingly overlap with the final response as reasoning progresses. We introduce SSR, a training-free self-speculative decoding method that uses a model's own partial-CoT response to accelerate generation of its final response, deriving both draft and verifier from the same model at different reasoning budgets. The draft is generated concurrently with the remaining reasoning, hiding most of its cost, and is reused through prefix verification and suffix decoding to speculate the final response. SSR accelerates the response rather than reasoning and requires no auxiliary models. Across three model families (Gemma 4, Qwen3, and GPT-OSS), SSR achieves 1.23–1.55× speedups over standard decoding on long-form code generation (ClassEval) with comparable task accuracy; the draft alone yields 1.22–1.27×, and on top of suffix decoding with the same lookup mechanism, SSR adds a further 6–8% on Gemma and Qwen. We also analyze when SSR helps, including the effect of draft timing and length and where the reused text comes from. Code will be released upon publication.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.