TreeSpark: From Acceptance Scores to Useful Work in Conditional Speculative Decoding
Abstract
Speculative decoding accelerates autoregressive language models by using a small drafter to propose tokens that a larger target verifies in parallel. Diffusion-based drafters can produce a whole block in one forward pass, and lightweight conditional heads can model dependencies within that block. Extending these proposals to trees covers more possible continuations, but the extra construction and verification work can outweigh the benefit of accepting more tokens. We study this trade-off with TreeSpark, a conditional-tree decoding framework built on the DSpark drafter. TreeSpark preserves a primary chain and adds optional leaves using already computed parent distributions and the drafter’s existing confidence head. A threshold chosen from the active batch size controls leaf admission, while ordered rejection verifies the sampled tree. We evaluate this design against chain and fixed-tree controls in the same optimized inference engine. On Qwen3-4B, selected fixed trees improve throughput over chain by 5.0–8.6%; confidence-guided trees gain 5.8–6.7% in sampling and 2.4–3.3% under fixed-rate serving. Matched static and constant-threshold controls do not establish a consistent additional benefit from adaptation. The results show why useful tree allocation must account for executed work, alongside acceptance predictions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.