acceptodds
Under review as a conference paper at ICLR 2027

Distribution Matching Distillation for Continuous Diffusion Language Models

Abstract

Continuous diffusion language models are becoming competitive with their discrete counterparts, but generation still requires many sequential network evaluations. Trajectory-based distillation methods have been developed to reduce this cost, while distributional distillation remains underexplored in this setting, particularly in exploiting the discrete structure of language and enabling multi-step generation. We develop a unified formulation of existing distributional distillation approaches that clarifies how the student’s probabilistic parameterization determines the available gradient estimators. Building on this formulation, we introduce two methods that share the same student architecture and reverse-KL objective: Simplex-DMD uses a continuous token relaxation and pathwise gradients, while Reinforce-DMD uses categorical token sampling and REINFORCE with a learned density ratio. We extend both methods to multi-step generation and investigate the training and sampling choices associated with each parameterization. On OpenWebText, using a sequence length of 1,024 tokens, Simplex-DMD and Reinforce-DMD improve the generative perplexity–entropy frontier over competing methods at complementary sampling budgets. In particular, at comparable entropy, Simplex-DMD reduces generative perplexity by 49% at 4 network evaluations, while Reinforce-DMD achieves a 20% reduction at 256 evaluations, each relative to the best evaluated baseline at the same budget.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.