A-Flow: Perturbation Can Be Correction for Frozen Flow Planners in End-to-End Driving
Abstract
Recent end-to-end driving methods use generative models for trajectory planning. They apply reinforcement learning (RL) fine-tuning to improve driving safety and comfort. However, such fine-tuning can be costly and sensitive to training data. Parameter updates may also disrupt pretrained driving behavior. Inspired by non-reversible Langevin dynamics, we design a lightweight skew-symmetric matrix, denoted by , to parameterize a structured perturbation that is learned as a trajectory correction. Building on this design, we propose A-Flow, a driving framework with three components. First, Temporal BEV Mapping (TBM) combines camera history to represent road geometry and traffic dynamics for trajectory generation. Second, we use GRPO to search for high-reward trajectories around the uncorrected trajectory. We optimize only while keeping the pretrained planner frozen. The learned matrices form a correction library for reuse. Third, Field-Guided Correction Transfer (FCT) retrieves relevant matrices to generate corrected candidates in unseen scenes. It uses TBM features to evaluate these candidates and select the final trajectory. A-Flow achieves 90.7 PDMS in NAVSIM v1 and an HD-Score of 37.49 in HUGSIM. HUGSIM evaluation uses zero-shot transfer without target-scene reward optimization. Experiments on a ReCogDrive-based planner further demonstrate the adaptability of our correction method.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.