Don’t Rewrite, But Correct: Closed-Loop Manifold Correction via RL Post-Training for Generative Driving Policies
Abstract
Generative driving policies have become a dominant source of pretrained driving priors in autonomous driving. In particular, flow-based planners trained by imitation learning capture expert behavior well on routine scenarios, whereas a few long-tail states can trigger compounding errors during closed-loop interaction, leading to a substantial gap between offline generative quality and closed-loop performance. To this end, we propose Neighborhood Action Correction, a closed-loop neighborhood post-training framework for generative driving planners. Different from the typical RL fine-tuning approach that updates the whole pretrained planner and may reshape its behavior on routine states, NAC freezes the generative planner and only trains a lightweight corrector to adaptively correct its prior action, particularly on long-tail states. The trainable correction is bounded by a neighborhood and confined to the only variable that enters vehicle dynamics under receding-horizon execution. NAC enables efficient improvement of the planner’s closed-loop performance while preserving the strong prior manifold of the generative policy. Experiments on nuPlan show that NAC achieves state-of-the-art closed-loop performance with flow-matching planner and outperforms the evaluated RL fine-tuning methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.