World-Language-Action Models for Autonomous Driving
Abstract
Autonomous driving requires understanding the observed scene, predicting the consequences of actions, and planning safe trajectories. Existing efforts to combine these capabilities often use separate visual representations for understanding and generation. We propose Doe-3, a world-language-action model that jointly models world states, language, and actions. We adopt the pretrained driving world representation to encode synchronized multi-view observations into scene tokens that retain appearance, semantics, and geometry. We introduce a shared autoregressive backbone that treats this world representation as a modality alongside language and action. The backbone provides a common temporal context for grounded language responses, trajectory planning, and action-conditioned future prediction. We attach a lightweight diffusion head that maps autoregressive world tokens to future scene latents, which pretrained decoders convert into RGB, depth, semantic, and occupancy outputs. We jointly train the backbone with language, planning, and future-prediction objectives to connect semantic understanding with scene dynamics. We adopt a three-stage schedule that progressively aligns the scene interface, adapts the backbone, and jointly optimizes the model. Doe-3 achieves a PDMS of 94.0 on NAVSIMv1 and an EPDMS of 90.8 on NAVSIMv2, outperforming the strongest prior methods. On nuScenes, it reduces the future-depth absolute relative error to 0.13 and reaches 27.03 mIoU and 61.49 IoU in camera-based 4D occupancy forecasting.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.