GeoWAM: Visual Geometry World Action Models for Autonomous Driving
Abstract
World action models (WAMs) have recently gained increasing attention as a framework for jointly modeling scene evolution and ego actions in autonomous driving. Most existing WAMs learn scene dynamics in pixel space by combining a video-generation backbone for future-observation prediction with an action head for ego-trajectory prediction. Pixels, however, provide only an indirect representation of these dynamics: they entangle geometry and motion with appearance, texture, and illumination, forcing the model to infer three-dimensional transformations from two-dimensional observations. We argue that point-based geometry provides a more natural state space for driving. It explicitly captures spatial structure and both rigid and non-rigid scene dynamics while remaining aligned with the 3D space of driving actions. Building on this insight, we introduce GeoWAM, a visual geometry world action model for autonomous driving. Rather than predicting future images, GeoWAM is pretrained to forecast future scene geometry, yielding representations that jointly encode spatial structure and temporal evolution. A geometry-conditioned action head then leverages these learned geometric dynamics to predict future ego-trajectories. Extensive experiments show that GeoWAM outperforms image-based alternatives, achieving an EPDMS of 90.2 on NAVSIM v2 and a combined EPDMS of 36.6 on navhard without PDMS supervision. Scaling geometry pretraining raises the navhard score by 8.2% to 39.6 and enables strong zero-shot transfer to nuScenes, yielding an average L2 error of 0.89 m and a collision rate of 0.12%. Together, these results establish geometry as an effective state representation and geometry pretraining as a general, scalable strategy for downstream planning.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.