Training expressive diffusion world models from scratch in under 50 GPU hours
Abstract
Autoregressive video diffusion has emerged as a promising approach for training expressive world models which can simulate complex interactions in dynamic environments. However, existing video diffusion models are typically highly expensive to train, which can often limit open-source research to fine-tuning of existing models or inference-time approaches with limited flexibility. In this work, we introduce WDDQ, a computationally efficient training recipe for a world model for Minecraft. In particular, we develop an autoencoder with a highly compressed latent space, a high-fidelity causal diffusion decoder, and a world model architecture combining an efficient masked predictor with a diffusion model that generates latent frames conditional on predictor outputs. Using this recipe, we train a world model that outperforms existing open-source models while training from scratch using less than 50 GPU hours. Unlike prior work, we accomplish this without distilling pretrained visual encoders, overfitting on specific scenes, or leveraging privileged information about 3D positions or geometry.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.