acceptodds
Under review as a conference paper at ICLR 2027

Cross-View World Models

Abstract

World models enable agents to plan by imagining future states, but existing approaches operate from a single viewpoint, typically egocentric, even when other perspectives would make planning easier. We introduce Cross-View World Models (XVWM), trained with a cross-view prediction objective: given frames from one viewpoint, predict the future state from the same or a different viewpoint after an action is taken. Because the input and output views may share little or no visual overlap, the model must learn view-invariant representations of the environment's 3D structure and the agent's place in it. Trained on synchronized multi-view gameplay data from Aimlabs, the model localizes itself on an overhead map from egocentric frames alone and renders egocentric scenes from only a small marker on that map. Probing the diffusion transformer layer by layer pinpoints where these representations emerge and how actions transform them. In early layers, the encoding of the agent's position and heading depends on the input view; in the middle layers it converges onto a view-invariant code, where a linear probe trained on egocentric inputs decodes pose from bird's-eye inputs without retraining, and vice versa. Within this code, heading lies on a ring distributed across many units. When we inject a yaw action, the ring rotates rigidly by the commanded angle, while the heading encoded in the earlier, view-dependent layers is left untouched: the model applies the turn to its view-invariant code, not to the view-specific encoding that precedes it. Together, these results give a mechanistic account of the model's “cognitive map”: how it tracks where it is and which way it is facing in a known environment. Such mechanistic understanding of how world models represent an agent's spatial state, and how actions transform it, is essential for reliable embodied agents.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.