Token Freezing in Diffusion Language Models Needs No Convergence Test
Abstract
Masked diffusion language models recompute every position at every denoising step, including tokens that are already committed and will not change. Freezing a committed token saves this compute by serving its cached keys and values, at the cost of staleness for the positions that still read them. The prevailing rule freezes a token once a convergence statistic of its output distribution, such as stepwise KL divergence, falls below a threshold. We show that this statistic cannot see the damage freezing does. Since a committed token is never rewritten, freezing affects the output only through the keys and values that other positions read. The statistic, however, measures the token's own output distribution. At a committed position, this distribution already places nearly all of its probability on the committed token, so it barely changes even when the keys and values do. To test whether the statistic's timing matters, we randomly shuffle its freeze times across tokens without increasing total compute. On every model we test, accuracy does not measurably change. We therefore drop per-token signals and freeze every token a fixed number of steps after commitment. We show that, among schedules that use no per-token information, this uniform delay minimizes the expected motion that frozen keys and values have left. Across five open diffusion language models, the delay matches or exceeds the prevailing rule at equal compute and improves on it by two to three points on held-out MATH. Rules built on other per-token signals, including the motion of the keys and values themselves, do not beat it either. The delay's savings carry over to wall-clock time, with 1.3–1.4× speedups at batch size 8 and above.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.