Understanding and Enabling Score Distillation in Continuous Diffusion Language Models
Abstract
Score Distillation (e.g., Diff-Instruct and Trajectory Distribution Matching) has substantially accelerated image and video diffusion models, yet transferring it to continuous diffusion language models requires more than reusing the training objective. We identify a central distinction in the relationship between denoising predictions and generated samples: in categorically parameterized language models, continuous distillation updates act on expected token embeddings, whereas generation ultimately requires discrete token decisions. Consequently, increasing prediction confidence can either improve generation or collapse different samples onto a small set of tokens. We propose a unified perspective: effective language distillation should resolve token-level uncertainty while preserving noise-dependent variation across generated samples. A denoiser's conditional mean should therefore not be assumed to provide a suitable generator initialization. This perspective unifies one-step endpoint matching and few-step trajectory matching, guiding generator initialization and score modeling of the student's generated distribution. Building on these principles, we adapt DI and TDM to continuous diffusion language models, enabling effective one-step and few-step generation on LM1B and OpenWebText. Our work clarifies the key considerations that distinguish distribution-matching distillation for language and provides reusable principles for transferring image-diffusion distillation methods to language generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.