acceptodds
Under review as a conference paper at ICLR 2027

Carryover Drafting: Recycling Rejected States for Speculative Decoding

Abstract

Speculative decoding accelerates LLM inference by verifying multiple drafted tokens in parallel, allowing a single target forward pass to generate several tokens. By construction, verification computes representations for both accepted and rejected tokens, but conventional drafters retain only the representations of accepted tokens, leaving the substantial verifier computation spent on rejected tokens effectively wasted. We show that these discarded hidden states retain useful information about future tokens that can improve subsequent drafts. However, realizing this opportunity poses two distinct challenges. At inference, reusing rejected states introduces additional projection and attention costs, which can increase drafting latency and offset the speedup gained from higher acceptance. During training, standard parallel training does not expose the drafter to rejected states encountered at inference, while generating such states through sequential rollouts would sacrifice parallelism across training positions. We introduce Carryover Drafting, which enables efficient reuse of rejected states and parallel inference-aligned training. Carryover recycles rejected target hidden states as temporary KV context for the next draft, allowing the drafter to selectively attend to the rejected states through its attention layers. A single learned embedding distinguishes rejected states from committed context, while the carried-over KV context is replaced each drafting round to bound the additional context by one proposal block. We further introduce parallel draft–verify–draft training to expose the drafter to inference-aligned rejected states without sacrificing parallelism across training positions. Across DFlash and a DSpark-derived semi-autoregressive drafter, Carryover improves average acceptance length by up to 14.7% and end-to-end serving speedup by up to 14.4% over the corresponding baselines, with gains reaching 28.8% on translation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.