acceptodds
Under review as a conference paper at ICLR 2027

Measuring and Reducing the Train–Inference Mismatch in Discrete Diffusion Models

Abstract

In discrete diffusion, we train a denoiser model on states from a noising process (i.e., on corrupted data). At inference time, the model has to reverse this process, iteratively denoising data, and is thus exposed to states created by its own samples. This creates a train–inference mismatch: the discrepancy between the distribution over states presented to the model during training and inference. We formalize this mismatch as the Mirror Gap. We measure this gap using three measures of discrepancy: the discriminability of training and inference states, the distance between their representation centroids, and the divergence between their quantized representation distributions. Together, these measurements give us a window into how the mismatch behaves at different resolutions and across the diffusion process. Our analyses reveal the train–inference mismatch has a low-dimensional component in representation space; they also show correlations between discriminability scores early in sampling and final sample quality. These results motivate two strategies to reduce this gap: steering generations in representation space, and resampling states based on measured gap. Experimentally, we show these strategies improve the generation compute–quality curve: at the same quality (measured by MAUVE or FID), our methods show sampling-efficiency gains exceeding 8× over strong baselines while requiring no denoiser retraining, fine-tuning, or distillation; this is true in both text and image generation, and across noising processes and samplers.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.