Where Draft Trees Lose Target Mass: Exit-Guided Speculative Decoding
Abstract
Tree-based speculative decoding organizes draft tokens into a bounded tree so that multiple continuations can be verified in one target-model pass. However, the tree is constructed from draft-side scores while its usefulness is ultimately determined by the target model, creating a fundamental draft–target mismatch under a finite tree budget. We study two questions: once a draft tree is fixed, can a better exact verifier accept more draft tokens, and if not, how can target feedback improve the tree itself? We answer both through a target-flow view of the fixed draft tree. We identify a canonical exit law that describes where target continuations leave the tree, and prove that one plus the resulting target coverage is a sharp upper bound on the expected output-block length, including the bonus token, of any exact path verifier. Moreover, every verifier attaining this ceiling must realize the same exit and bonus-token law; representative predraw-and-follow and sequential residual verifiers already attain it. This theory directly yields Tree Exit Verification (TEV), an exact verifier that realizes the canonical law through one exit-node decision and one bonus-token decision, exposing a regular, level-parallel verification procedure. The same exit law also localizes where target probability is missing from the current tree. We use it as node-level feedback to train the drafter on its inference-time draft trees, reallocating finite tree budget toward target-relevant regions. Experiments across dialogue, code, and mathematical reasoning validate the fixed-tree equivalence, show that the Exit-Guided Draft-Tree Training (ExitTrain) improves average output-block length by 13%, and that TEV reduces verifier-stage latency by 15%, yielding 14% end-to-end speedup over DDTree. Together, our results separate the two remaining opportunities in bounded tree speculative decoding: better draft trees for higher acceptance, and more direct verification for lower latency. Code is available at https://anonymous.4open.science/r/TEV-ICLR-348C/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.