acceptodds
Under review as a conference paper at ICLR 2027

Looking Ahead: Modeling Future Evolution in Autoregressive Video Generation

Abstract

Autoregressive video diffusion models enable video generation by sequentially denoising video chunks conditioned on preceding context. Beyond producing successive chunks, a key capability expected from this paradigm is to model how visual scenes evolve over time, which is also central to world modeling. However, teacher forcing, together with the strong temporal redundancy of video, makes current-chunk denoising susceptible to shortcut learning, allowing the model to exploit local visual continuity without sufficiently capturing the dynamics of video evolution. To bridge this gap, we introduce Future Chunk Prediction (FCP), which augments autoregressive training with future-chunk supervision that is less directly supported by local visual continuity, encouraging the model to better capture how visual scenes evolve over time. We evaluate physical plausibility as a challenging manifestation of video evolution modeling. Our method achieves superior physical plausibility over standard teacher-forced autoregressive training with a lower overall training compute budget, highlighting the effectiveness and efficiency of future-chunk supervision for modeling visual evolution. Intriguingly, FCP also induces more spatiotemporally structured representations, suggesting a promising direction toward predictive representations for world modeling.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.