DriveBeyond: Rethinking Group Advantage for Driving Policy Optimization
Abstract
End-to-end autonomous driving models are typically initialized through imitation learning (IL), after which reinforcement fine-tuning (RFT) is increasingly used to improve driving quality. Group Relative Policy Optimization (GRPO) has been adopted for this stage, estimating each trajectory's advantage from its reward relative to other trajectories sampled for the same scene. We revisit this group-relative construction and characterize two issues in its advantage estimation under driving RFT: advantage shift, where centering within each group discards the scene-level progress over training, and advantage collapse, where candidates receiving identical rewards yield vanishing advantages. Motivated by these findings, we propose DriveBeyond, a driving policy optimization method that incorporates references beyond the current sampling group through a shared memory module, calibrating advantages for groups with sufficient reward variation and helping recover learning signals for groups without. Experiments on NAVSIM v1 and v2 demonstrate that DriveBeyond consistently improves planning performance across both VLA- and WAM-based planners under three evaluation settings.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.