acceptodds
Under review as a conference paper at ICLR 2027

When Better Drafts Cost More Verification: Safety Boundaries for Multi-Token Drafting

Abstract

Speculative-decoding draft heads are trained to match the target model at each position. Yet a better local draft can make a whole response more expensive to verify: accepting more tokens moves where verification restarts, possibly to a costlier continuation. We characterize this effect for a fixed target path with a fixed proposal cap, full-prefix commitment and unit cost per verifier call. First, same-rank readouts can raise some target-match probabilities, lower none, yet raise verifier calls in proportion to the cap. Second, every such improvement is safe when the expected accepted length is at most one at every restart, the largest uniform threshold, or when prefix-survival probabilities are consistent across restarts; a sharp bound prices departures from consistency. Third, an exact backward recursion decides safety outside both regions, certifying 370 of 10,752 fitted probability fields against every coordinatewise increase in match probabilities, where the threshold alone certifies two. Finally, in a preregistered ten-seed comparison on two frozen backbones with matched anchors and evaluation budgets, whole-response training wins none of eight comparisons against a local objective covering every anchor, and the local objective reduces sampled verification calls in four. Draft-head changes should therefore be judged by whole-response verification work; our conditions and exact test show when a local improvement is guaranteed safe.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.