Training Reshapes Width-Dependent Acceptance in Block-Diffusion Speculative Decoding
Abstract
Inference-time width selection balances a drafter's acceptance against execution cost. We study how training changes that balance. In block-diffusion drafting, a narrower forward can lower acceptance even when the verified prefix is unchanged. A three-arm experiment separates training forward width from the positions included in the loss. With identical prefix supervision, narrow-layout training improves acceptance at narrow evaluation widths and lowers it at a wider width in all three paired seeds. On 900 prompts from six sources, the gain extends to width 8, the configuration selected for serving. This finding motivates cost-driven adaptation: identify promising training layouts from measured costs, adapt with standard cross-entropy, then select each head's deployment width independently. On MI308X, both supervision-matched heads select width 8 from the same candidates. Narrow-layout training improves throughput by 1.58% on average, with positive gains in all three seeds and sessions and 17 of 18 paired comparisons during nine-minute runs. Public-checkpoint and cross-platform tests show how the benefit varies across models and workloads. Training thus changes the acceptance–cost trade-off available to a deployment policy.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.