From Chains to Trees: Parent-Conditioned Drafting for Semi-Autoregressive Speculative Decoding
Abstract
Speculative decoding accelerates LLM inference only when drafted continuations survive target-model verification. Semi-autoregressive drafters such as DSpark predict a token block with one backbone forward and refine it with a lightweight Markov head. However, DSpark decodes this block as a chain, so an early mismatch invalidates the remaining suffix and limits the benefit of large draft blocks. We show that the conditional structure already learned by DSpark can support multiple parent-consistent continuations without retraining or additional backbone passes. We introduce Parent-Conditioned Drafting Tree (PCTree), which uses the pretrained Markov head to score alternative children separately for each parent and allocates a fixed verification budget to the most probable paths. This converts DSpark's linear draft into a tree while preserving its one-pass parallel backbone. Across Qwen3-{4B,8B,14B} and nine benchmarks, at , measured speedup gains over autoregressive (AR) decoding, relative to matched DSpark, range from % to %. On Qwen3-4B GSM8K at , \method increases mean acceptance length from to and three-run mean AR speedup from to . These show that parent-conditioned branching can turn conditional capacity already present in a semi-autoregressive drafter into end-to-end inference gains through an inference-only change.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.