Data-Forcing Distillation: Restoring Diversity and Fidelity in Few-Step Video Generation
Abstract
Recent progress has shown promise in distilling multi-step video diffusion models into efficient few-step students. Among them, Distribution Matching Distillation (DMD) and its successor DMD2 achieve strong generation quality and fast convergence. However, due to the nature of the reverse Kullback–Leibler (KL) objective, these methods can exhibit two persistent failure modes: reduced sample diversity and visibly over-saturated outputs that deviate from real-video appearance. In this work, we propose *Data-Forcing Distillation* (DFD), a simple post-training framework that improves DMD-based students *with only a single line of code change*. At its core is the *teacher score discrepancy*, which guides the student toward the reference-data distribution by pulling probability mass toward underrepresented modes while suppressing problematic modes that are inconsistent with the reference data. We provide an in-depth theoretical analysis of our framework and evaluate it on text-to-video, image-to-video, and autoregressive video generation. With only *100–300 steps* of finetuning, DFD improves both sample diversity and visual quality in text-to-video and autoregressive generation, while improving visual quality and temporal coherence in image-to-video generation. These results show that DFD can mitigate key failure modes of few-step video generators across different generation settings. We refer readers to the supplementary material for full video comparisons.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.