On-Policy Forcing: Autoregressive Video Generation without Bidirectional Distillation
Abstract
Recent autoregressive (AR) video generators often rely on a separate bidirectional teacher for few-step distillation, limiting the role of AR training as a native foundation for the full generation pipeline. In this work, we show that teacher-forcing AR diffusion training can produce strong many-step video generators. We then introduce **On-Policy Forcing**, which directly distills the trained AR model into a few-step student while preserving the same causal factorization from AR training through few-step distillation. By using AR models for both real- and fake-score estimation during on-policy distillation, the method removes the need for an additional bidirectional teacher and the conditioning mismatch of asymmetric distillation, in which bidirectional score models see future context that the AR student cannot. With all methods evaluated by our VBench pipeline, our four-step student distilled with 1.3B AR score models reaches 84.08 Total, outperforming Self Forcing and Causal Forcing distilled with 1.3B bidirectional score models by 0.86 and 0.79 points, respectively. At an equal training budget, On-Policy Forcing improves Dynamic Degree on the official VBench subset by 28.33 points over distillation with bidirectional real- and fake-score models, while a conditioning mismatch in the fake-score model alone leads to severe degradation. These results validate AR training as a viable foundation for video generation and establish a fully causal training-to-distillation recipe for AR video models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.