One Schedule Does Not Fit All: Adaptive Inference Schedules for Diffusion Language Models
Abstract
Diffusion language models commonly rely on a single, fixed denoising schedule across all workloads, even though the cheapest quality-preserving allocation of model evaluations varies substantially across tasks and checkpoints. We introduce Task-Calibrated Denoising Allocation (TCDA), a confidence-controlled calibration framework that selects the lowest-cost schedule from a structured candidate library using a paired calibration sample and a user-specified quality tolerance. TCDA compares candidate topologies, including early prefix decoding, state-reusing online continuation, reduced full-range resampling, and the canonical baseline, deploying the cheapest schedule certified by one-sided simultaneous confidence bounds and task-specific degeneration guardrails. Across three frozen ELF-B tasks, TCDA identifies three distinct operating regimes: early prefix stopping for XSum, state-reusing online continuation for WMT, and guarded full-range resampling for OpenWebText. At representative tolerances ( for XSum and WMT; 20% relative perplexity for OpenWebText), selected schedules reduce denoiser evaluations by to , yielding measured wall-clock speedups between and with held-out violation rates below . Under strict margins, the method safely defaults to the full schedule. These findings establish that optimal inference reduction is workload-dependent, providing a disciplined methodology for budget-constrained diffusion inference.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.