Beyond the Tip of the Iceberg: Depth-Aware Top- Margin Distillation for Diffusion Speculative Decoding
Abstract
Speculative decoding accelerates LLM inference by using a lightweight drafter to propose future tokens for parallel verification. Diffusion-based and block-parallel drafters can generate long speculative sequences efficiently, but their acceptance often degrades rapidly with speculative depth, particularly at longer context lengths. We introduce Depth-Aware Top-K Margin Distillation, a training objective designed to improve drafter–verifier alignment by accounting for the unequal importance of speculative positions. At each depth, our objective distills the verifier's distribution over its Top-K candidate tokens and aligns the logit margins between adjacent ranked candidates, while assigning larger weights to earlier positions whose rejection invalidates the remaining speculative suffix. The formulation requires no modification to the drafter architecture or inference procedure and can, in principle, be applied to different speculative drafter families, while being particularly well suited to diffusion and parallel drafting where depth-wise acceptance degradation is pronounced. We apply our objective to a block-diffusion drafter with three target models (Qwen3-8B, Qwen3-14B, and Gemma-4-31B) on five benchmarks. With Qwen3-8B and Gemma-4-31B, it improves the average acceptance length by 15–16% over the same drafter trained with its original objective. With Qwen3-14B, it improves acceptance length by 28% over a published baseline that uses a smaller proposal budget. These gains come with correspondingly higher end-to-end speedup. The benefit grows with context length: on BABILong, the relative gain in acceptance length rises from 28% at 8K tokens to 55% at 1M tokens. These gains come with no additional inference-time overhead.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.