Normalizing Flows are Capable World Models
Abstract
World models predict how the world evolves in response to actions and have developed along two main directions. Video world models, most of which are built on diffusion models, generate future observations with impressive quality but typically require large models and multiple denoising steps per prediction. Latent world models instead predict and plan in a learned representation space. Those based on Joint Embedding Predictive Architectures (JEPAs) learn this space jointly with the predictor through a regression objective, but require an additional anti- collapse mechanism, and their deterministic predictors cannot represent multiple possible futures. To bridge these directions, we present Normalizing-Flow World Model (NFWM), an autoregressive normalizing flow consisting of an invertible encoder and a probabilistic predictor. Both components are trained jointly by exact maximum likelihood, which trains the predictor to model a distribution over future latent states and prevents representation collapse without an additional anti-collapse mechanism. Since the encoder is invertible, NFWM can generate future observations like a video world model, while predicting in a single step and planning in its learned latent space like a latent world model. Although the encoder preserves all information, its latent coordinates are still learned, so we can shape them for planning with an auxiliary inverse-dynamics objective. Empirically, NFWM achieves the highest average success on deterministic goal-conditioned planning and outperforms JEPA-based world models in stochastic environments and on robotic manipulation, while also supporting policy learning in imagination. These results provide initial evidence that normalizing flows are capable world models, establishing them as a promising direction for connecting video and latent world models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.