acceptodds
Under review as a conference paper at ICLR 2027

Can Video World Models Track Unobserved World States?

Abstract

Video world models are increasingly used as simulators, but visual fidelity alone does not show that a model maintains the hidden state of the world. We examine this difference with an action-conditioned video Shell Game, a visual analogue of state tracking that separates visual rendering from compositing the unobserved world state. Trained on 5-swap chains, standard backbones (e.g., bidirectional and autoregressive Transformers, Mamba, and linear attention) render plausible videos and predict the correct ball location up to 5 swaps. However, they fail to learn the rule and generalize to longer swap chains, even with more denoising steps. As the pixel-based diffusion loss does not force the generated frames to hold the unseen ball position, output tokens cannot carry it, and the state has to live within the architecture. In a causal Transformer, this implicit state is an append-only KV cache, which is written once and never revised, so the model must re-compose the swaps at every chunk. Tracking this way requires depth to grow with sequence length, which no fixed-depth Transformer provides. We study what enables learning the rule, and find that length generalization requires a revisable state carried across chunks and an update expressive enough to apply a swap. Linear attention can achieve this by allowing negative transition eigenvalues, and autoregressive Transformers can do so with nonlinear TTT fast weights (e.g., SwiGLU) whose online updates change the feature map used to read their state. We further examine Memory Maze and Block World, where the state is not fixed by the input action stream alone and must be corrected from observations or keeps changing out of view, and discuss the implications for building stateful video world models.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.