acceptodds
Under review as a conference paper at ICLR 2027

Video World Models Need Predictable Latents

Abstract

Video world models trained on large-scale robot data promise a safe and scalable approach for learning and evaluating policies that act in the physical world. However, they still struggle to generate realistic long-horizon rollouts, generalize out of distribution, and they remain computationally demanding at inference time. The dominant approach, making predictions in the latent space of a reconstruction-trained autoencoder, does not incentivize a latent geometry that facilitates dynamics modeling. For this reason, we investigate whether encouraging a meaningful temporal organization in latent space can benefit downstream world model performance, using a series of controlled experiments. We find that temporal representation learning objectives can improve downstream performance of the world model at matched reconstruction quality. However, the best-performing objectives need over 20x more autoencoder updates to reach the same reconstruction performance. To address this issue, we characterize the trade-off between information retention and useful geometric organization in latent space training, and we propose a simple modification to the architecture of the decoder that improves it. We find that pairing this architecture with temporal objectives for autoencoder finetuning improves downstream performance of the world model at matched training budget. Compared to a reconstruction baseline, the gains include up to 5x faster inference, 8x faster training, 24% MSE decrease out of distribution and 18% MSE decrease in long-horizon extrapolation. This also translates into better agreement of the world model with real-world policy evaluation, with up to a 0.46 increase in Pearson correlation using 2-step inference. Common to the best temporal objectives is a large increase in one-step predictability of the latents, lower intrinsic dimensionality, and weaker high frequency components. Overall, our study indicates that temporal predictability is an important design criterion for the latent space of video world models, and provides concrete guidance on how to incorporate it into autoencoder training.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.