DIAL: One Diffusion Language Model, Multiple Inference-Time Compute Budgets
Abstract
Block diffusion language models trade generation quality for inference cost by varying the number of tokens committed per denoising step, yet this block size is typically fixed during training, requiring separate specialized models for different deployment settings. We present DIAL, a training framework for block-size-elastic diffusion language models in which the inference block size is a runtime choice shared by a single parameter set. The core mechanism, mixed block-size training, samples block sizes throughout training, exposing one network to a range of denoising granularities from autoregressive-like factorization to highly parallel generation. To preserve performance across this range, DIAL combines per-block noise conditioning with a zero-initialized recurrent draft-refinement loop, a two-stream cached parameterization, and argmax-agreement early stopping. Experiments show that mixed block-size training substantially reduces the degradation that otherwise occurs when a model is evaluated away from its training block size, while a single DIAL model achieves competitive perplexity across all tested inference granularities. By treating block size as an inference-time control, dial exposes multiple quality–compute operating points from one trained model, allowing deployment to adapt generation cost to the available compute budget without retraining or maintaining separate block-size specialists.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.