acceptodds
Under review as a conference paper at ICLR 2027

Calibrate Before You Branch: Feedback-Calibrated, Cost-Aware Trees for Block-Diffusion Speculative Decoding

Abstract

Speculative decoding (SD) accelerates large language model inference by using a lightweight drafter to propose future tokens that are verified in parallel by the target model. Block-diffusion speculative decoding further reduces drafting latency by predicting multiple future positions in a single forward pass. To exploit alternative candidates within such parallel drafts, recent methods organize them into draft trees for joint verification. However, existing tree construction primarily ranks candidate paths using factorized draft probabilities, which may not reflect the target model’s actual acceptance preferences. We propose FACT, a cost-aware draft-tree allocation framework that jointly addresses candidate prioritization and budget selection. Specifically, FACT uses historical target-verification outcomes to construct a frozen correction table that calibrates candidate priorities without additional inference-time model calls. It further combines draft-coverage estimates with profiled full-round costs to adaptively determine whether expanding the tree is worthwhile, while an incremental builder reuses the search frontier across budget checkpoints and materializes only the selected tree. Experiments on six benchmarks with Qwen3-8B and Qwen3-4B show that FACT consistently increases tokens per verification round over DDTree-B128. On Qwen3-8B, FACT achieves 4.17–6.80× speedup over autoregressive decoding and reduces generation time by 0.27–2.73% relative to DDTree-B128. On Qwen3-4B, it reduces generation time by 1.24–3.05%. Code is available at https://anonymous.4open.science/r/fact-anonymous-review-B853/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.