DVGT-2: Vision-Geometry-Action Model for Autonomous Driving
Abstract
Vision-Language-Action (VLA) and World-Action (WA) models offer promising routes to end-to-end autonomous driving. In this paper, we introduce an alternative Vision-Geometry-Action (VGA) paradigm and demonstrate that dense geometric supervision can support competitive driving planning. As vehicles operate in a 3D world, dense geometry provides rich spatial cues for decision-making. However, the computational overhead of multi-frame geometry reconstruction limits its use in online driving. To address this, we introduce a streaming Driving Visual Geometry Transformer (DVGT-2), which jointly predicts local geometry and future trajectories, aligning geometric reconstruction with ego-centric planning. We employ temporal causal attention and cache historical features to avoid repetitive computation. We further propose a sliding-window streaming strategy and use relative temporal encoding to support online inference with bounded per-frame computation and memory. Experiments demonstrate superior local geometric accuracy across five driving datasets and substantially improved online reconstruction efficiency. DVGT-2 achieves strong planning performance on NAVSIM v1/v2 and remains competitive on nuScenes, while offering low inference latency compared with representative VLA and WA models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.