Separating marginal error from recoverable dependence in parallel-head speculative decoding for large language models
Abstract
Speculative decoding commits several tokens per forward pass by checking a block a cheap drafter proposed. When that drafter is parallel heads, one per future position, they propose independently and miss the target's joint law for two confounded reasons: marginal error, and dependence no product can represent. We separate them exactly at depth two for two deployed head families whose first proposal is the target's own next-token head. Under that condition the deficit is second-marginal error plus a recoverable term, the total variation lost when the target's conditional mismatch is marginalised over the first position. The target's own dependence bounds that term, giving a normalised recoverable fraction that a certified estimator measures at vocabulary scale. On a fixed target, marginal error strongly predicts that fraction without determining it: at fixed prefixes the certified bounds establish that proposals with exactly equal second-marginal error can differ in recoverable fraction, and the target's own head does not behave like the perturbations that concentrate mass. Across three seeded runs on a frozen trunk, the calibration differs between prefix distributions of different first-position entropy, and a positive entropy-residual coupling appears with training. Because the recoverable term caps what dependence can add at fixed marginals, a small value rules out a large gain from that redesign before it is built.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.