acceptodds
Under review as a conference paper at ICLR 2027

Lost in the Rollout: Rethinking On-Policy Distillation for Multimodal Video Diffusion

Abstract

On-policy distillation, exemplified by Self Forcing, has become the leading approach for turning video diffusion models into fast, interactive few-step generators. Applying it under joint text, image, and audio conditioning for avatar generation has recently attracted great interest, but rare insights have been exposed on the corresponding training recipe. This paper first identifies the lost in the rollout issue in this regime, where the video rollouts generated by the student often lose what they should carry at two levels. First, they lose the conditions: within a few hundred steps, training collapses into black frames, flickering, or blur, and the identity given by the image condition is gone. Unlike a text prompt, the image condition becomes the first frame of every rollout, and the teacher and the critic score each rollout under that same image, so a flaw in the image enters the very signal meant to correct the student. (1) Curated conditions and (2) a converged ODE initialization keep the conditions intact and stabilize training; together they cut FID from 27.10 to 11.67 and raise Sync-C from 3.13 to 4.15. Second, they lose the teacher's trajectory: after the ODE initialization, the student sees only the DMD score of its own rollouts, and at standard learning rates and guidance it learns too little lip-sync. We keep a small piece of the regression onto teacher samples that DMD2 removed: (3) an auxiliary target that matches the audio energy of each generated clip to that of a full teacher sample for the same condition. Because the target is computed on the generated audio, it applies to any model that generates sound. (4) Higher learning rates and teacher guidance make the student learn more from the teacher's score, raising Sync-C further to 4.50. With these remedies, a collapsing recipe becomes a stable four-step causal student of a Wan2.1-based multimodal model, which improves lip-sync, image quality, and aesthetics over the bidirectional model of the same size on HDTF, AVSpeech, and CelebV-HQ, with over 20× higher throughput, running at 24.82 FPS on a single GPU. On MiniMax-H3, a 33B model that generates video and audio jointly, the auxiliary target further sharpens lip-sync, raising Sync-C from 7.37 to 8.02. The lesson reaches beyond these models: an on-policy student excels when its rollouts start well and the teacher's signal keeps reaching them.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.