Closing the Loop on Open-Loop Driving Data: Reinforcement Learning of Action Policies with Controllable World Models
Abstract
Policies trained from recorded driving logs rarely observe how their own decisions change future observations, limiting their ability to recover once they depart from the recorded trajectory. We present DriveVerse, an autoregressive multi-view world model that turns logged scenes into interactive environments for policy learning. At each decision step, a structured simulation backend executes part of the policy plan and constructs view-aligned map and motion conditions; DriveVerse then generates synchronized camera observations for the next decision. These post-action observations enable recurrent policy interaction beyond the recorded ego trajectory. Shared cross-view attention, causal training, and self-rollout distillation support recurrent generation, while online diagnostics monitor rollout reliability. We optimize the resulting rollouts with T-GRPO, combining plan-level and executed-trajectory feedback across vision-language-action, world-action, and end-to-end policies through their native action representations. Experiments show that DriveVerse outperforms the strongest baseline, lowering front-view FVD by 22% and improving all reported object- and lane-geometry metrics and temporal consistency. Interactive training with DriveVerse improves all three policy families (VLA, WAM, E2E) over their SFT and RL baselines on both navtest and navhard. The VLA reaches 92.8 and 52.3 EPDMS, respectively, improving on its GRPO baseline by 1.7 and 5.2 points.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.