acceptodds
Under review as a conference paper at ICLR 2027

Temporal Context Needs the Right Target: Projected Latent Prediction for End-to-End Driving

Abstract

Temporal context is essential for motion-aware driving, yet we find that naively introducing temporally fused features into a reconstruction-based latent world model can degrade planning performance. We attribute this failure to a representation–objective mismatch: the predicted latent encodes temporal history and action-conditioned evolution, whereas its supervision target remains a single-frame visual latent whose exact coordinates contain factors unnecessary for planning. To address this mismatch, we propose Temporal-JEPA LAW. Our method first introduces Latent World-Model Temporal Fusion (LWM-TF), which causally aggregates each LAW latent token over time while preserving its spatial identity and maintaining attention complexity. It then replaces raw latent reconstruction with normalized projected alignment, allowing the world model to match prediction-relevant structure without reproducing every visual-latent coordinate. Following LeWorldModel, a sketched isotropic Gaussian regularizer (SIGReg) prevents collapse in the unnormalized embedding space. Controlled experiments reveal the central effect of this design: temporal fusion alone improves average planning L2 from 0.61m to 0.53m, whereas coupling temporal prediction with raw mean-squared error degrades it to 0.66m. Projected alignment mitigates this negative transfer, achieving 0.51m average L2 and reducing the average collision rate from 0.30% to 0.09%. These results show that temporal information and predictive supervision must be designed jointly: LWM-TF supplies motion evidence, while JEPA provides a planning-compatible target for latent future prediction.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.