ST-JEPA: JEPA-Enhanced Temporal Scene–Trajectory Representation Learning for End-to-End Autonomous Driving
Abstract
End-to-end autonomous driving requires compact representations that preserve planning-relevant spatial and temporal context. A common design in recent ViT-based E2E planners is to introduce trajectory queries only after visual feature extraction, which limits the interaction between planning intent and visual representation learning during encoding. Meanwhile, in compact-token variants, scene tokens receive only auxiliary supervision without direct self-supervision on their latent representations. To address these limitations, we present **ST-JEPA**, a temporal scene–trajectory representation learning framework with two key designs: scene–trajectory encoding and JEPA-enhanced representation learning. First, ST-JEPA injects camera-aware scene tokens and ego-conditioned trajectory tokens directly into the visual encoder, enabling them to interact with image patches throughout feature extraction. Interleaved temporal global-register attention further aggregates cross-frame context and propagates the updated representations through subsequent backbone blocks. Second, ST-JEPA incorporates joint-embedding predictive learning through two complementary mechanisms. An online scene–trajectory JEPA objective predicts compact scene and trajectory representations from masked observations, using targets produced by an EMA teacher from unmasked inputs. This training-only objective provides latent self-supervision directly at the compact planning bottleneck. In parallel, we adapt JEPA pretraining to driving videos and integrate the resulting spatiotemporal latent features into the encoder by projecting them into the token space and fusing them with the native scene representation. ST-JEPA achieves state-of-the-art performance on NAVSIM-v1, NAVSIM-v2, and the closed-loop HUGSIM benchmark, with systematic ablations validating the effectiveness of the proposed scene–trajectory encoding, temporal interaction, and JEPA-based representation learning. The code will be made available upon publication.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.