acceptodds
Under review as a conference paper at ICLR 2027

DreamerAD: Efficient Reinforcement Learning via Latent World Model for Autonomous Driving

Abstract

Training reinforcement learning (RL) policies through real-world interaction is costly and risky. Video-generation-based world models can provide realistic video rollouts in place of real-world interaction, but multi-step denoising and RGB decoding make these imagined rollouts computationally expensive. We introduce DreamerAD, a framework for efficient reinforcement learning through latent-space imagination for autonomous driving. Our approach leverages denoised latent features from video generation models through three key mechanisms: (1) shortcut forcing that distills multi-step flow updates into low-step latent rollouts, accelerating world model inference by up to , (2) an autoregressive dense reward model that predicts step-wise rewards directly from latent representations, providing temporally dense feedback for policy optimization, and (3) Gaussian vocabulary sampling for GRPO that constrains exploration to physically plausible trajectories. DreamerAD achieves 91.4 EPDMS on NAVSIM v2 and 55.0 HD-Score on zero-shot HUGSIM, establishing state-of-the-art performance across both benchmarks. Across both the driving-specific Epona world model and the general-purpose Wan2.2 video generation backbone, reinforcement learning through latent-space imagination improves zero-shot closed-loop driving performance on HUGSIM.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.