acceptodds
Under review as a conference paper at ICLR 2027

Mass Forcing: On-Policy Mass-Aware Distillation for Autoregressive Video Generation

Abstract

Distribution Matching Distillation (DMD) has emerged as a promising paradigm for real-time video synthesis, distilling pretrained diffusion models into few-step generators. However, its reverse-KL objective is mode-seeking and can collapse the student's conditional video distribution onto a narrow subset of modes, leading to reduced sample diversity, diminished temporal dynamics, and color oversaturation. Although existing methods primarily mitigate this issue via coverage-oriented student initialization, we observe that such initial coverage is transient: across three different initializations, the fraction of dynamic videos drops sharply during subsequent DMD training, indicating that the unchanged mode-seeking objective, rather than the initialization, drives this degradation. Motivated by this training-time degradation, we propose Mass Forcing, a probability-mass-aware framework that incorporates coverage-oriented supervision throughout on-policy DMD training via complementary local and set-level objectives. First, at the local level, Mass Forcing estimates relative teacher-to-student density ratios across rollouts sharing a student-generated prefix through latent-space score integration and uses these estimates to rebalance conditional DMD updates, amplifying the training influence of rollouts underrepresented relative to the teacher. Second, recognizing that mismatched mode probabilities do not necessarily induce strong local score corrections, Mass Forcing aligns generated rollout sets with a reweighted reference distribution defined by the same mass estimates, providing cross-rollout supervision beyond individual gradient reweighting. Experiments on MovieGen and VBench demonstrate that Mass Forcing improves video fidelity and sustains temporal dynamics throughout training, achieving the highest Total scores among compared methods and a Dynamic Degree of 83% after 1,000 generator updates, versus 12-32% for the baselines, and introducing no additional inference-time computational overhead.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.