Frequencies Rank, Probabilities Allocate: Calibrated Draft Trees for Drafter-Free Speculative Decoding
Abstract
Drafter-free speculative decoding verifies trees of continuations retrieved from the prompt, the output so far, or a datastore, and allocates each tree by retrieval scores, such as n-gram frequencies, that rank candidates. We show that ranking is not enough: over all candidate trees, best-first allocation is invariant to a continuous, strictly increasing transformation of edge scores if and only if it is a power map; scores that order every set of siblings correctly can still lose a factor of K in expected accepted draft tokens; and errors of conditional acceptance probabilities along root paths bound the loss. Motivated by source- and depth-dependent miscalibration of retrieval frequencies on real traces, we introduce CalTree, a 29-parameter acceptance model fitted on annotation-free replays of the target's own generations, over retrieval and request-local evidence. On five targets from 1.5B to 32B parameters, it generates 2.6–6.5% more tokens per verification pass than SSSD on Spec-Bench and 4.5–24.5% more on code editing, at every tested budget and target. Source and depth discounts tuned for this metric at each budget on the same replays recover 24–53% of the Spec-Bench gain, and normalizing CalTree's probabilities over the retrieved siblings keeps their order but forfeits 1.3–5.0 points at the tested budgets. Scores fitted on chat transfer to code, and a Qwen2.5-7B scorer to four other targets, without refitting. In fixed-load SGLang benchmarks, CalTree-LR raises throughput over SSSD by 1.0–3.0% on Spec-Bench and 2.6–4.0% on code editing; full-workload outputs are not always token-identical to autoregressive decoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.