Collective Ego-Agent World Modeling for End-to-End Autonomous Driving
Abstract
World modeling enables end-to-end autonomous driving (E2EAD) to anticipate scene evolution and learn from action consequences. However, existing methods predominantly adopt an ego-centric perspective, treating surrounding agents as scene content rather than explicit rollout subjects. This formulation limits the completeness of future scene representations and leaves predictive consistency across ego and agent perspectives underconstrained. In this paper, we propose CoDreamer, an E2EAD framework for collective ego-agent world dreaming. CoDreamer integrates an ego-centric branch and selected agent-centric branches through a shared latent dynamics model. To improve spatial completeness, each agent branch reconstructs future features over its full spatial window from a coordinate-transformed, partially observed Bird’s-Eye-View (BEV) representation, learning to infer how regions beyond its spatial support evolve. To encourage cross-perspective agreement, ego and agent rollouts are conditioned on their respective logged trajectories and jointly supervised at the prediction horizon by coordinate-aligned views of the same realized future. Ego-Agent Spatial Tokens (EAST) encode relative geometry and entity type, enabling the shared predictor to distinguish the rollout subject from other participants across coordinate frames. This collective learning objective provides additional supervision for planning representations, with future world-model transitions confined to training in the evaluated configuration. On NAVSIMv1, CoDreamer achieves state-of-the-art planning performance among world-model-based methods using the same ResNet-34 backbone, attaining 92.0% PDMS. Ablation studies further demonstrate improved prediction accuracy and scene completion beyond observed regions. These findings underscore the value of agent-centric world modeling for capturing more complete and coherent scene dynamics. Code will be released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.