GeoSCope: Grounded, Stable and Controllable Driving World Models
Abstract
Generative driving world models produce realistic video but lack grounding in a real location. Conditioned on their own predictions, they modify road layout and surrounding structures, causing long-horizon rollouts to diverge from the recorded scene. A policy evaluated in such rollouts is effectively tested on a different scene than the one recorded. We propose GeoSCope, a method that grounds driving world models in the geometry of a real route using only a single monocular recording. A dynamics-aware SLAM pipeline reconstructs a static 3D map of the route, which is rendered at the camera pose of each frame prior to its generation and used to condition a generative world model. Although the rendered signal is sparse and imperfect, the model learns to exploit it selectively. Experiments with three world models on four datasets show that this anchoring substantially reduces long-horizon drift and preserves scene structure without suppressing dynamic objects. Since the map can be rendered along arbitrary trajectories for conditioning, the same model supports counterfactual rollouts, turning existing video-based world models into grounded, stable, and controllable neural simulators. With geometric conditioning trained only on nuPlan, the approach transfers zero-shot to Waymo, MARS and nuScenes, and sets a new state of the art on the Ann-Arbor-City-Bench multi-traversal benchmark. Code and weights will be released upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.