RaPhy-AD: Radar-Aware Foundation Model with Explicit Physical Chain-of-Thought for Autonomous Driving
Abstract
Vision-language models (VLMs) have enabled autonomous driving systems to predict trajectories through intermediate reasoning steps. However, VLM-based reasoning is often driven by visual appearance rather than grounded in the physical quantities underlying the scene, such as the distance and relative motion of surrounding agents. Visually driven reasoning can generate linguistically plausible rationales that still remain inconsistent with the actual physical state of the scene, creating a physical-grounding gap between inferred reasoning and directly measured physical quantities. To bridge this gap, we propose RaPhy-AD, a radar-aware foundation model that extends VLM perception beyond visual by incorporating millimeter-wave radar and generates structured physical rationales grounded in radar measurements of object range, azimuth, and Doppler-based radial velocity. To represent these radar-derived physical quantities, RaPhy-AD integrates a Physical-Aware Tokenizer for numerical values and physical units and a Polar-Time Radar Encoder that preserves the temporal, azimuth, and range structure of multi-frame radar observations. We further develop a two-stage training strategy consisting of radar-aware supervised fine-tuning followed by Group Relative Policy Optimization with verifiable physical rewards to first learn radar-grounded physical knowledge and then refine the numerical accuracy of physical reasoning and trajectory prediction. RaPhy-AD shows strong trajectory prediction performance across challenging driving scenarios, particularly under visual degradation, highlighting the importance of radar-grounded physical information.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.