Learning to Drive from the Future, Planning Trajectories in One Forward Pass
Abstract
World–action models (WAMs) enrich autonomous-driving policies with future-state supervision, yet it remains unclear which visual representations best support joint future-state and action learning and how to retain the benefits of future supervision in a single-pass planner. We introduce JEWAM, a Joint-Embedding World–Action Model for joint future-state and action learning, and use this framework to evaluate different visual representations. Comparisons of features from ViT, MAE, CroCo v2, DINOv3, and V-JEPA under a shared interface show that representation choice substantially affects planning. V-JEPA's predictive video features achieve the strongest planning results among the evaluated representations, motivating their adoption in the final model. We further show that future-state supervision improves action learning while its benefits are retained with single-pass, action-only inference. Direct full-horizon prediction outperforms the tested iterative variants at lower inference cost, supporting effective world–action learning without future rollout or iterative denoising at deployment. Using only front-view camera observations, JEWAM achieves 90.98 EPDMS on NAVSIM v2 NavTest and 39.86 on NavHard.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.