acceptodds
Under review as a conference paper at ICLR 2027

Driving-OPD: Learning Complementary Driving Experts with Reward Variance-Aware RL and On-Policy Distillation

Abstract

Reinforcement learning (RL) is increasingly adopted for post-training of Vision- Language-Action (VLA) models in end-to-end autonomous driving. However, existing approaches mainly optimize a single policy, leaving the effect of differ- ent RL treatments on driving capabilities underexplored. Under structured driv- ing rewards, we find that GRPO’s variance-dependent scaling changes the rela- tive optimization emphasis across rollout groups. Motivated by this, we intro- duce reward-variance-aware RL, which explores different treatments of variance- dependent scaling to induce complementary driving specializations. Specifically, Dr.GRPO removes standard-deviation normalization and yields a safety-oriented expert, while our Variance-Gated Scaling GRPO (VGS-GRPO) retains controlled variance-dependent scaling and produces a progress-oriented expert from the same base policy and RL data. Building on this finding, we propose Driving- OPD, a three-stage post training framework consisting of supervised fine-tuning (SFT), reward-variance-aware RL for complementary expert learning, and On- Policy Distillation (OPD) for expert integration. We use Win-Margin-Based Sce- nario Selection and Dual-Teacher Routing to improve multi-expert distillation by providing more effective expert-specific supervision. Without explicit Chain-of- Thought reasoning or additional planning modules, Driving-OPD directly gener- ates trajectories autoregressively and achieves 92.8 PDMS on NAVSIM v1 and 91.4 EPDMS on NAVSIM v2, highlighting the strong post-training potential of autoregressive Driving VLAs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.