acceptodds
Under review as a conference paper at ICLR 2027

Motion-JEPA: Physically Structured Residual Predictive Representation Learning for Autonomous Driving

Abstract

Driving world models must predict and represent future scenes relevant to driving decisions. Changes in high-dimensional representations arise both from referenceframe changes caused by ego motion and from temporal evolution of the surrounding environment. When a model directly predicts the complete future, a predictor with limited capacity must fit both known geometric effects and unknown scene changes. This can impair future-scene prediction, produce representations that are less useful for driving decisions, and impose substantial computational costs. We propose MotionJEPA, a physically structured residual predictive representationlearning framework. It uses low-dimensional, interpretable ego-motion variables to construct a physical reference evolution for high-dimensional representations, and then learns the conditional residual not explained by that reference. The method combines explicit SE(2) motion transport, residual prediction around a physical reference, and static-dynamic diagnostics of residual sources. In 10,000 training steps on NAVSIM, model A3, which combines motion transport with an environment-residual predictor, reduces masked Smooth-L1 error by 7.76% relative to A1, which uses motion conditioning alone. A3 also achieves a 7.35% lower best validation Macro Smooth-L1 and a 7.69% lower overall test Smooth-L1 than A1. The additional validation metric, mean Intersection over Union (mIoU), also improves.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.