D-VOLT: Training Speculative Drafters with Token-Conditional Acceptance Values
Abstract
Speculative decoding accelerates language model inference by verifying several draft tokens in a single target-model forward pass. Training a lightweight drafter requires allocating its limited capacity to improve the expected number of accepted draft tokens. Existing approaches improve local distribution matching or reweight losses across positions, but candidate tokens at the same prefix can enable different expected numbers of subsequent acceptances. We identify a limitation of weighting a local acceptance objective by a detached continuation weight shared across candidate tokens at the same prefix and position. Even when the weight uses the exact mean continuation value conditional on acceptance, the resulting objective can favor updates that increase one-step acceptance probability while decreasing expected accepted length. We propose D-VOLT (Drafting via Value-Optimized Lookahead Training), which assigns training credit using token-conditional continuation values. We derive a Bellman recursion over complete prefix histories and an exact gradient decomposition of expected accepted length, in which each candidate's credit combines prefix reachability with its immediate acceptance reward and continuation value. For practical training, we approximate continuation values by dynamic programming over a compact top- candidate graph that shares sampled earlier histories across branches. We combine full-vocabulary total-variation matching with a candidate-specific continuation reward, retaining the drafter architecture and standard verification procedure. Experiments with DSpark drafters on BBH, MMLU-Pro, WinoGrande, and GSM8K at decoding temperatures and show that D-VOLT improves average accepted length by 4.05%–9.74% over the DSpark baseline and by 2.04%–6.46% over D-PACE.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.