DualWM: Dual-Layer Unified Asynchronous Latent World Model for End-to-End Autonomous Driving
Abstract
Latent world models enhance autonomous driving by learning task-state transitions, providing valuable policy priors. However, existing methods suffer from the fundamental conflict between cognitive reasoning and real-time responsiveness in dynamic driving scenarios. To address this challenge, we propose a novel end-to-end autonomous driving framework based on a Dual-layer Unified Asynchronous Latent World Model (DualWM) that learns state transitions at both the intent and behavioral levels within a unified latent space for efficient driving. Specifically, the VLM-based upper layer handles scene understanding and long-horizon intent, while the lower layer enables real-time semantic state prediction. The two layers are mutually coupled: the intent features provide long-horizon priors that guide behavioral prediction, and the behavioral states in turn constrain the uncertainty inherent in VLM predictions, while the asynchronous architecture improves inference efficiency. This design draws inspiration from human driving behavior, where conscious planning coexists with rapid reactive adjustments. Finally, a Mixture-of-Experts (MoE)-based policy module is designed to capture both the consistency and conflicts between macroscopic intent and behavioral-level decisions, enabling efficient and robust closed-loop driving. Extensive experiments on the Bench2Drive closed-loop benchmark in CARLA demonstrate that DualWM achieves strong closed-loop driving performance while substantially improving runtime efficiency, validating its potential for real-world autonomous driving.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.