acceptodds
Under review as a conference paper at ICLR 2027

Beyond Imitation: Planner-Agnostic Reinforcement Learning Post-Training for End-to-End Autonomous Driving

Abstract

End-to-end autonomous driving has advanced rapidly on large-scale closed-loop benchmarks such as nuPlan, with planners commonly based on regression, diffusion, or flow matching. Although imitation learning (IL) provides strong behavioral priors, it primarily optimizes similarity to logged expert trajectories rather than task-level driving metrics such as safety, progress, comfort, and interaction quality. Reinforcement learning (RL) can address this limitation, but existing RL post-training methods are often coupled to planner-specific generation mechanisms, resulting in fragmented solutions across planner families. We propose Planner-Agnostic Reinforcement Learning (PARL), a unified post-training framework that isolates planner-specific differences in a lightweight log-probability adapter while sharing the remaining RL pipeline, including reward computation, KL regularization, and policy optimization. PARL supports regression-based, diffusion-based, and flow-matching-based planners without modifying their architectures, inference interfaces, or action spaces, and requires no privileged information. Under the full nuPlan closed-loop protocol, covering non-reactive and reactive settings on Val14, Test14-hard, and Test14, PARL consistently improves all six evaluated IL planners across all six settings, with the largest gains in interaction-heavy long-tail scenarios. These results demonstrate the feasibility and effectiveness of unified, planner-agnostic RL post-training for end-to-end autonomous driving. Code and pretrained checkpoints will be released upon acceptance.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.