Mixture-of-Latents for Few-Step Discrete Diffusion Distillation
Abstract
Discrete diffusion and flow models can generate high-quality sequences, but they usually require many denoising steps. Distilling them into few-step models is challenging because a single student step predicts many positions independently, which can lose important dependencies between tokens. We show that matching each position separately is not enough to reproduce the teacher's joint transition. Prior work addresses this issue with a shared continuous latent variable, but its training requires an approximation that can fail to make effective use of the latent. We propose Mixture-of-Latents for Discrete Diffusion Distillation (MLD3), which instead uses a small finite set of components selected by a learned router. This makes the training objective exactly computable and allows the student to model correlated updates directly. We also propose a novel sampling approach that exploits the mixture components at a cost close to that of a factorized step. Notably, MLD3 naturally reduces to a standard masked diffusion student with one component and adds negligible parameter overhead as new components are added. Across text, molecules, regulatory DNA, and image generation, MLD3 consistently improves few-step generation by roughly 2x over the baselines, with the largest gains on tasks with stronger dependencies across positions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.