Implicit EOS Density for Online Bidirectional Length Adaptation in Masked Diffusion LLMs
Abstract
Masked diffusion LLMs (e.g., LLaDA) typically rely on a fixed generation length (i.e., a fixed MASK canvas), struggling to accommodate responses whose required lengths vary substantially across inputs: a short length may constrain content generation, while a long length wastes computation. We study the denoising dynamics and find that the evolving implicit density () of end-of-sequence (EOS) tokens during denoising reveals whether the current length is excessive or insufficient. Building on this observation, we propose -EOS, a training-free, online bidirectional length adaptation strategy for masked diffusion LLMs: low density triggers expansion, while high density triggers contraction. This enables online length adaptation within the original denoising loop, without retraining or a separate length-adjustment stage. Extensive experiments on mathematical reasoning and code generation tasks show that -EOS yields a more favorable Pareto frontier in the quality-efficiency trade-off. Additionally, because -EOS targets online length adaptation, it is orthogonal to pruning, caching, and other acceleration-oriented decoding strategies, making it potentially compatible with and complementary to them.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.