acceptodds
Under review as a conference paper at ICLR 2027

DriveAnchor: Progressive Anchor-based Flow Learning for Autonomous Driving Planning

Abstract

Autonomous-driving planners must generate route- and maneuver-compliant trajectories under a limited decoder-call budget while adapting offline to black-box static-collision feedback. We propose DriveAnchor, a three-stage framework built around a shared vocabulary of recorded trajectories and a continuous trajectory interface. The vocabulary supports flow-matching (FM) residual learning, Energy Field (EF) corridor-conditioned initialization, and anchor-guided zeroth-order reward fine-tuning. The reward update operates in trajectory space without policy likelihoods or reward gradients, enabling offline adaptation of a deterministic decoder. On sampled subsets of nuPlan Val14 and Test14-hard per simulation mode, DriveAnchor achieves the highest reactive scores among the compared methods: 81.50 and 65.74, respectively. We further integrate the trajectory generator into a real vehicle's planning pipeline and demonstrate its operation on urban roads. On 10,000 held-out internal scenarios, under the original non-EF prior and a fixed two-call FM budget, reward fine-tuning reduces short- and long-horizon static candidate collision rates from 27.22% and 56.54% to 2.94% and 7.15%, respectively, and improves 8-s best-of-set minADE from 2.27 to 1.34 m. In a separate junction diagnostic on checkpoints before Stage 3, one and two EF calls raise corridor-guidance satisfaction from 22.10% to 74.13% and 87.92%, respectively. Additional ablations support the benefits of real-trajectory priors and anchor-guided exploration. Code and deployment videos are available at the anonymous project page: https://driveanchor.github.io/. In real-vehicle tests, trajectory generation takes 2.06 ms on average.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.