How to Train Your Video World Model for Long-Horizon Rollout
Abstract
Video world models must sustain autoregressive generation long after their initial observations leave the attention window. How should they be trained for this regime? Recent work advocates conditioning on clean history during pre-training to match inference, but the pretrained model is often only the starting point for further training. We study how this choice affects the complete pipeline for action conditioned world models without a separate bidirectional teacher. In this setting, the pretrained causal model provides both the student initialisation and the teacher for post-training. On Minecraft, we compare Diffusion Forcing, clean prefix teacher forcing, and chunked teacher forcing at equal pre-training compute. Each model undergoes flow map finetuning followed by Self Forcing against its own teacher, with evaluation over 1000 generated frames. The recipe rankings change after post-training. Diffusion Forcing initially performs worst beyond the attention window, but subsequently outperforms both alternatives through frame 400 and matches chunked teacher forcing at longer horizons. Its Fréchet Video Distance over frames 401–1000 falls from 329 to 86, compared with 132 to 91 for chunked teacher forcing. Clean prefix teacher forcing is best at neither stage. The post-training stages play complementary roles: qualitative rollouts suggest that flow map finetuning increases dynamic content and Self Forcing improves its persistence. These results show that alignment with inference during pre-training alone does not determine rollout quality. Training recipes should be evaluated through the full pipeline and at long horizons.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.