acceptodds
Under review as a conference paper at ICLR 2027

DVGT-World: Extending Vision-Geometry Action Paradigms To 3D World Models For Autonomous Driving

Abstract

Autonomous driving requires not only understanding the current 3D scene but also predicting how it will evolve. Recent Vision-Geometry-Action (VGA) models have advanced end-to-end driving through explicit 3D geometry reconstruction. However, most VGA methods focus on reconstructing currently observed scenes, lacking explicit modeling of future 3D dynamics. While World-Action Models (WAMs) forecast scene evolution, their predictions in 2D image or latent spaces lack explicit 3D geometric supervision. To address this, we present DVGT-World, a framework that extends the VGA paradigm into a predictive 3D world model. By jointly supervising current geometry reconstruction and future geometry prediction, DVGT-World learns representations that capture both spatial structure and scene evolution, providing a geometric foundation for downstream planning. Specifically, we introduce learnable future queries conditioned on the target horizon into the geometry transformer to predict future dense pointmaps and ego poses from historical multi-view observations. After geometry pretraining, we fine-tune the model for planning, with future prediction disabled for inference efficiency. Experiments demonstrate accurate multi-view future geometry prediction and strong closed-loop planning performance, achieving 93.3 PDMS on NAVSIM v1 and 90.7 EPDMS on NAVSIM v2, which validates the benefit of modeling 3D scene evolution through future geometry prediction for downstream planning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.