acceptodds
Under review as a conference paper at ICLR 2027

ACTOR-JEPA: MULTI-AGENT MOTION FORECASTING USING LATENT-PREDICTIVE PRETRAINING

Abstract

Motion forecasting models must capture how agents influence one another, and are trained on large labeled trajectory datasets. Self-supervised pretraining promises to reduce that dependence on labels, yet reported results conflict, ranging from substantial gains to null results. We introduce Actor-JEPA, which applies the Joint-Embedding Predictive Architecture to multi-agent scene encoding. Its encoder factorizes the scene into a cross-agent interaction stage followed by per-agent temporal encoding, and that factorization defines the prediction target: agents are mixed before masking, so the model predicts masked agent-timestep positions in a representation that already carries interaction context, rather than reconstructing coordinates. Across four label regimes and independent pretraining runs on the Waymo Open Motion Dataset, Actor-JEPA raises peak mAP from to at labels ( relative, , ) and from to at full labels (, , ), in each case against the identical architecture trained from scratch. Within a dataset, two conditions govern whether any benefit appears: the cross-agent stage must be transferred, since the per-agent core alone performs at from-scratch level, and pretraining must stop before a representation collapse that neither the pretext loss nor the variance regularizer reveals, which we detect by monitoring the effective rank of the encoder's features. Code will be made publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.