World Models in Robotics and Autonomous Driving: A Survey of the Latency-Fidelity Trade-off and the Case for Unified Scene-Motion Dynamics
Abstract
World models – learned predictive models that compress an agent's observations into a latent representation and forecast how that representation evolves under action – have become a unifying substrate for both simulated control and autonomous driving. This paper synthesizes thirty such systems (2018–2026) along four axes: the five training objectives in current use (variational latent dynamics, discrete autoregressive token prediction, denoising score matching, joint-embedding feature prediction, and value-equivalent dynamics); four architecture families (recurrent/variational, Transformer/diffusion, Joint-Embedding Predictive Architectures, and driving/robotics-specific systems); the choice of 3D scene representation, from raw pixels to bird's-eye-view latents, occupancy grids, and point clouds; and five largely incommensurable evaluation protocols, under which only about a quarter of latency-relevant papers report any throughput figure at all. A persistent tension between representational richness and inference latency runs through every family: the most expressive generative world models are, by their own authors' account, not yet real-time, while real-time-capable models sacrifice visual fidelity, multi-agent modeling, or explicit 3D grounding. We argue that unifying scene and motion representation within a single, low-latency model – rather than treating world modeling and action planning as separate stages – is the central open challenge for physical, real-time robotic and autonomous-driving deployment, and examine three candidate directions toward it – continuous-time (Neural ODE) dynamics, cross-embodiment generalization, and structured state-space backbones – each grounded in evidence from the surveyed corpus or its margins, though none yet demonstrated at the scale the gap requires.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.