Prefix-Conditioned Diffusion: Reducing the Pretraining–Generation Mismatch in Diffusion Language Models
Abstract
Diffusion language models (dLLMs) generate text through iterative denoising, allowing multiple tokens to be predicted in parallel. However, pretraining may mask tokens throughout a sequence, whereas prompt continuation conditions on an intact prefix. This difference remains in conversion pipelines that denoise entire sequences during the stable stage. We propose Prefix-Conditioned Diffusion (PCD), which samples a boundary, preserves the prefix, and denoises the suffix. The training recipe also applies autoregressive supervision to the prefix. We evaluate PCD in the stable stage of a warmup, stable, and decay conversion pipeline, with inference unchanged. Matched experiments across model families show improvements in reasoning and coding over native diffusion training. A matched continuation study further shows that the advantage persists after a shared decay stage. In a separate reconstruction diagnostic, the full PCD recipe's advantage over native diffusion training reverses as more evaluation prefix tokens are masked. Controlled experiments show lower suffix reconstruction loss with a clean training prefix, an intact evaluation prefix, and no autoregressive loss.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.