acceptodds
Under review as a conference paper at ICLR 2027

Correct While Verifying: Intermediate Target Predictions for Speculative Draft Correction

Abstract

Speculative decoding accelerates large language model inference, yet its rejection handling remains reactive. Even parallel approaches such as PEARL overlap drafting with verification, but the system still discovers that a draft token is wrong only after final verification and discards the entire suffix after the first rejection. We make a simple observation: target verification is layered, and intermediate target computation already produces useful predictions about the correct continuation before final logits are available. Based on this observation, we correct drafts while they are being verified. A trainable decoder layer maps intermediate target states to preliminary token distributions, whose top- candidates seed correction branches. While the target model completes final verification from cached intermediate states, a frozen draft model expands these branches in parallel under a custom attention mask. Final target verification evaluates both the original draft and the corrected branches and commits the longest valid continuation under the target model's verification rule. Across Llama, Qwen2.5, and Vicuna target-draft pairs, our method consistently outperforms PEARL and achieves up to 4.53 speedup over autoregressive decoding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.