When Better Drafts Cost More Verification: Safety Boundaries for Multi-Token Drafting
Abstract
Speculative-decoding draft heads are trained to match the target model at each position. Yet a better local draft can make a whole response more expensive to verify: accepting more tokens moves where verification restarts, possibly to a costlier continuation. We characterize this effect for a fixed target path with a fixed proposal cap, full-prefix commitment and unit cost per verifier call. First, same-rank readouts can raise some target-match probabilities, lower none, yet raise verifier calls in proportion to the cap. Second, every such improvement is safe when the expected accepted length is at most one at every restart, the largest uniform threshold, or when prefix-survival probabilities are consistent across restarts; a sharp bound prices departures from consistency. Third, an exact backward recursion decides safety outside both regions, certifying 370 of 10,752 fitted probability fields against every coordinatewise increase in match probabilities, where the threshold alone certifies two. Finally, in a preregistered ten-seed comparison on two frozen backbones with matched anchors and evaluation budgets, whole-response training wins none of eight comparisons against a local objective covering every anchor, and the local objective reduces sampled verification calls in four. Draft-head changes should therefore be judged by whole-response verification work; our conditions and exact test show when a local improvement is guaranteed safe.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.