AD-E2E-JEPA: A Joint-Embedding Predictive Architecture For End-to-End Autonomous Driving
Abstract
Autonomous driving requires that can understand the physical world, reason and plan, and operate safely. In this paper, we first systematically evaluate existing action-conditioned joint-embedding predictive architecture (JEPA) world models, including LeWM, DINO-WM, and JEPA-WM for end-to-end autonomous driving (E2EAD). To isolate world-model quality from policy learning, we employ a goal-conditioned zero-shot planning setting that evaluates these models using ground-truth future observations as goals, without training any driving policy. We find that existing JEPA-based world models are either accurate for driving but computationally expensive, or computationally efficient but insufficient for planning. To address this trade-off, we propose , which introduces a SIGReg-regularized learnable projector applied to projected patch embeddings. The projector reduces the number of planning patches by and the embedding dimension by , achieving a inference speedup while retaining planning performance, with a 0.8-second runtime for an 8-frame rollout over 256 candidate trajectories. Without training any driving policy, the world model itself reaches the goals located 20 meters away on average within the displacement of respectively 4.0/3.0 meters, using world-model rollouts over trajectory vocabularies of respectively 256/2,048 candidates. On the NAVSIMv2 benchmark, it achieves 67.3/71.5 EPDMS with multiplicative safety metrics and 84.1/86.0 EPDMS without them in goal-conditioned zero-shot planning. Experiments further show that the self-supervised pretrained projector improves downstream imitation learning performance from 80.2 to 85.4 EPDMS.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.