acceptodds
Under review as a conference paper at ICLR 2027

Latent Prediction Helps Physical Representations Only When It Predicts the Future

Abstract

Joint-embedding predictive architectures (JEPAs), which predict in a learned latent space, are reported to learn better representations of physical systems than masked autoencoders, which reconstruct pixels. Previous studies, however, change the space of the prediction target together with its time horizon, the architecture and the training pipeline making it difficult to determine which factor contributes to improved representations. With this motivation, we separate prediction space from prediction horizon. We train ViT-S encoders on three physical simulations from The Well. These encoders are identical except for their objectives. We train on four distinct objectives, varying whether they predict (a) pixels or latent representations and (b) the current eight-frame clip or the next eight-frame clip. Across all encoder layers, we probe representations for regime parameters and for pooled and token-level physical quantities at current and future times. We find that regime probes and future-time probes favour representations trained on the Latent-Future objective, suggesting that latent prediction helps physical representations only when paired with future prediction. Probes for same-time physical quantities provide similar results for all models. We further analyse representations after adding noise to the inputs and find there is not one objective that produces representations more robust to noising. This paper analyses how prediction space and prediction horizon affect learned representations while holding the other training components fixed.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.