UBTree: Parallel Tree Drafting via Unigram and Bigram Models for Speculative Decoding
Abstract
Speculative decoding accelerates language model inference by verifying multiple draft tokens predicted by lightweight draft models in a single target-model forward pass. Recent parallel drafters have achieved breakthrough performance in frontier production models, yet their gains diminish as the target distribution becomes more entropic, particularly for later positions within draft blocks and under long-context settings, where limited draft diversity increasingly constrains acceptance. To overcome this bottleneck without sacrificing parallelism, we introduce **UBTree**, a parallel drafter that couples a **U**nigram proposer with a **B**igram selector to construct drafting **Tree**s. The unigram proposer is trained with the standard cross-entropy objective to generate candidate tokens independently for each position, while a lightweight bigram selector predicts transition scores between adjacent candidate pairs. Unlike the proposer, the selector is trained with a renormalized KL objective on high-temperature data. This *tree-native* training broadens the supervision beyond the greedy path, encouraging plausible alternative branches that improve the chance of accepting additional tokens during tree verification. Across seven standardized benchmarks with Qwen3-4B and Qwen3-8B, UBTree achieves an average speedup of – over autoregressive decoding and outperforms DARTree in all 28 comparisons. Production-scale evaluation further demonstrates UBTree's advantage over frontier baselines such as DSpark.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.