acceptodds
Under review as a conference paper at ICLR 2027

Distilling Uniform Discrete Diffusion Models into a One-Step Video Generator

Abstract

Discrete diffusion has emerged as a competitive paradigm for visual generation, yet high-quality generation still relies on repeated refinement steps, limiting its use in low-latency interactive video generation. Compressing this process into a single forward pass is particularly challenging for video, where tens of thousands of categorical variables must jointly establish coherent spatial structure and temporal dynamics without subsequent correction. In this work, we present DDMD-V, to our knowledge the first one-step distillation framework for uniform discrete diffusion video generation. DDMD-V extends on-policy distribution matching to the uniform discrete diffusion setting, using student-induced intermediate states along the teacher’s metric-dependent corruption path. To complement token-level distribution matching, we further introduce spatiotemporal adversarial regularization in the FSQ embedding space, providing a learned signal over multi-token spatial and temporal configurations. Distilled from URSA-1.7B, our student generates a 49-frame \(320\times512\) video comprising discrete tokens in a single network forward pass. It achieves 81.78 VBench score compared with the 81.75 score of the 50-step teacher, while reducing the number of function evaluations (NFEs) from 100 to one when counting the teacher's conditional and unconditional guidance branches. After distillation, the student achieves closely matched visual quality compared to the teacher. In contrast, directly reducing the teacher to one sampling step yields only 58.56 VBench. With BF16 inference, DDMD-V reaches a generation throughput of more than 50 video frames per second on a single H200 GPU. These results demonstrate that a strong discrete video diffusion model can be distilled into a one-step generator with comparable performance, providing a promising path toward low-latency generative video models. %for interactive world-model systems. We will release the code and model.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.