acceptodds
Under review as a conference paper at ICLR 2027

Calibrating Noise Schedules for Continuous Language Diffusion with Token Uncertainty

Abstract

Continuous language diffusion models use noise schedules to allocate training capacity and guide inference across varying corruption levels. Existing schedules often adapt heuristics from image diffusion, rely on costly denoiser queries, or assume uniform token distributions, overlooking how embedding geometry and token frequency affect recoverability under Gaussian noise. We present **kNN Information Scheduling** (KIS), an approach that bases noise scheduling on token frequencies and distances between embeddings without querying a denoiser. By modeling token uncertainty under Gaussian noise, KIS tracks conditional entropy across log signal to noise ratios (log-SNRs) and expresses its decrease as a normalized entropy reduction. Instead of replacing schedules without constraints, KIS applies ***constrained calibration***: it shifts the training distribution's median while keeping its standard deviation fixed and applies a positive affine correction to the interior sampling levels in log-SNR. Across various architectures (ELF, FLM) and datasets, KIS significantly accelerates inference and enhances generation quality. On FLM-B, KIS improves generative perplexity by up to 25.6% on OpenWebText and 47.1% on LM1B under limited budgets. Notably, a calibrated FLM at 256 steps outperforms its uncalibrated version at 1024 steps, using four times fewer sampling steps. Ablation studies confirm that affine calibration outperforms direct quantile placement at 8–32 steps in the reported comparisons, offering key insights into the relationship between discrete priors, continuous representations, and diffusion dynamics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.