acceptodds
Under review as a conference paper at ICLR 2027

Beyond Pixels: Grounding Video World Models with State-Centric Information

Abstract

Video world models are trained under the paradigm of predicting future video frames from a stream of input actions, with the hope of unlocking fast, interactive and physics-accurate simulation environments from video data only. Many state-of-the-art world models rely on off-the-shelf VAE encoders of text-to-video generative models, thus inheriting from their ad-hoc latent space, which is also hard to interpret or control, and is not built with the task of world modelling in mind. In contrast, in settings where hard-coded physics simulations are standard (e.g. autonomous driving or robotics), world models often rely on explicit 3D, sensor information or dedicated logic engines. In this work, we study how privileged state information can be used alongside video data to unlock new properties of video world models: We first show that a modular training recipe enables improved efficiency, while modeling both the states and video streams jointly induces a natural measure of uncertainty of the full world state, which may only be partially observable in the video. We also exploit the structure of the state space to infuse LLM priors during training as a form of data augmentation to correct policy bias or introduce novel actions. Finally, we design state-based metrics for the Rocket Science dataset to evaluate the world model's physics understanding.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.