Driving-OPD: Learning Complementary Driving Experts with Reward Variance-Aware RL and On-Policy Distillation
Abstract
Reinforcement learning (RL) is increasingly adopted for post-training of Vision- Language-Action (VLA) models in end-to-end autonomous driving. However, existing approaches mainly optimize a single policy, leaving the effect of differ- ent RL treatments on driving capabilities underexplored. Under structured driv- ing rewards, we find that GRPO’s variance-dependent scaling changes the rela- tive optimization emphasis across rollout groups. Motivated by this, we intro- duce reward-variance-aware RL, which explores different treatments of variance- dependent scaling to induce complementary driving specializations. Specifically, Dr.GRPO removes standard-deviation normalization and yields a safety-oriented expert, while our Variance-Gated Scaling GRPO (VGS-GRPO) retains controlled variance-dependent scaling and produces a progress-oriented expert from the same base policy and RL data. Building on this finding, we propose Driving- OPD, a three-stage post training framework consisting of supervised fine-tuning (SFT), reward-variance-aware RL for complementary expert learning, and On- Policy Distillation (OPD) for expert integration. We use Win-Margin-Based Sce- nario Selection and Dual-Teacher Routing to improve multi-expert distillation by providing more effective expert-specific supervision. Without explicit Chain-of- Thought reasoning or additional planning modules, Driving-OPD directly gener- ates trajectories autoregressively and achieves 92.8 PDMS on NAVSIM v1 and 91.4 EPDMS on NAVSIM v2, highlighting the strong post-training potential of autoregressive Driving VLAs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.