acceptodds
Under review as a conference paper at ICLR 2027

GEOVLDRIVE: GEOMETRY-GROUNDED VISION-LANGUAGE REASONING FOR AUTONOMOUS DRIVING

Abstract

Autonomous driving involves continuous interaction between the ego vehicle and surrounding agents in a dynamic 3D environment, requiring geometrically grounded reasoning about current spatial relationships and future scene evolution for safe and efficient planning. Despite advances in geometric grounding and future prediction, reasoning about the implications of evolving 3D spatial relationships for ego planning remains a fundamental challenge. We introduce GeoVLDrive, a geometry-grounded vision-language-action (VLA) framework for 3D spatiotemporal reasoning and planning. Multi-layer geometric features from a 3D foundation model complement the semantic representations of a visionlanguage backbone. A causal latent chain of thought (CoT) reasons from future scene geometry through key-agent motion to meta action. An action expert grounded in geometric reasoning jointly generates and scores trajectory proposals conditioned on these latent states and multi-depth scene features. We further introduce NAVSIM 3DQA to support training and evaluation of metric 3D spatial understanding and agent motion reasoning. GeoVLDrive outperforms the compared supervised VLA baselines on NAVSIM v1 with 91.8 PDMS, while attaining 90.9 EPDMS on NAVSIM v2. On NAVSIM 3DQA, it achieves an aggregate score of 89.2, outperforming a matched fine-tuned baseline by 4.6 points. It also achieves 68.6 on LingoQA.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.