FieldDrive: Field-Guided One-Step Trajectory Generation and Reward Optimization for Autonomous Driving
Abstract
Generative models have been introduced as trajectory decoders for autonomous driving to model the conditional multimodal distribution of reasonable driving behaviors in complex traffic scenes. However, existing generative decoders typically rely on multiple denoising or flow-integration steps, resulting in high inference overhead. Meanwhile, existing reinforcement learning methods are mainly designed around multi-step generation chains and are structurally mismatched with one-step generators. Combining one-step multimodal trajectory decoding with the optimization capability of reinforcement learning remains a central challenge in generative autonomous driving planning. To address this challenge, we propose FieldDrive, a Drift-based one-step trajectory decoder with a corresponding reinforcement learning method. We cluster a trajectory vocabulary and select high-scoring trajectories to construct a feasible trajectory distribution, which guides the generated trajectory distribution by the Drift Field Vector. We train a critic network to provide a Reward Field Vector that directs the model toward high-reward regions. Theoretically, we show that the Drift Field Vector and Reward Field Vector share a common origin: both arise as optimization directions of distributional energy function. Through a two-stage training paradigm, the model explores different distributions of reasonable trajectories and progressively converges to high-reward regions through reinforcement learning, while enabling efficient decoding with a single forward pass at inference time. Experiments on NAVSIM v1 and v2 demonstrate that FieldDrive enables efficient and high-quality trajectory decoding.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.