PhysDrive: A Physically Grounded Vision-Language Model for Driving Understanding and Planning
Abstract
Vision-Language-Action (VLA) models have emerged as a promising paradigm for autonomous driving, supporting both scene understanding and motion planning. However, existing VLA models are typically optimized primarily with downstream outputs such as language responses or expert trajectories, providing limited supervision for the intermediate representations that connect visual perception to decision making. This output-centric training may encourage shortcut learning and hinder the model from capturing the underlying 3D geometry and physical dynamics of driving scenes. To address this limitation, we propose PhysDrive, a physically grounded vision-language model that uses 4D occupancy as structured training-time supervision for its internal representations. Specifically, PhysDrive decomposes 4D occupancy supervision into two complementary components: the current occupancy grounds the model in the 3D geometry of the observed scene, while future occupancy guides the modeling of its physical evolution over time. Together, these signals encourage PhysDrive to encode spatial structure and scene dynamics before producing downstream understanding or planning outputs. We further introduce a reinforcement learning framework to align generated trajectories with driving-oriented objectives, including safety, drivable-area compliance, and progress. Extensive experiments on NuInstruct, nuScenes, and NAVSIM demonstrate that PhysDrive consistently improves driving understanding and planning, demonstrating the effectiveness of physically grounded representations for downstream driving tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.