Smarter Than Random, or Simply Busier? Where Diffusion-LLM Scheduling Earns
Abstract
Masked diffusion language models unmask several tokens per denoising step, and inference-time schedulers decide which positions to unmask. Every published scheduler is validated against a baseline that takes no action, so its gain cannot be told apart from the gain of simply acting more. We introduce the matched-action control, a random scheduler that takes exactly the same actions at randomly chosen positions, and run three rules against it across models, tasks and step counts. Under this control the gains separate into a substrate, the confidence ordering every scheduler shares, and an overlay, what each rule adds on top of it. The substrate is where the value lies, worth on open-ended text and to on conditional tasks on two models, while the overlay is exactly zero on open-ended text at every step count and appears on conditional tasks only at the most aggressive parallelism, on tasks that differ between models. The reason is a bound. Confidence ordering already maximizes the expected number of correct tokens under the model's own distributions, so a rule can add value only through dependence among the positions unmasked together, and that value cannot exceed the number of them that would change if they were unmasked one at a time. Measured directly, the bound is real and grows faster than linearly with the number of positions unmasked per step, yet the rules realize only 4% to 13% of it, because most of the changes it counts leave correctness untouched and the rest harm as often as they help. What choice is worth therefore lies in when rather than where. Unmasking few positions early and more later, with no rule and no extra cost, improves accuracy by up to on the same tasks, and the gain of the published method closest to our own turns out to be its schedule rather than its spacing.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.