Do-Causal-JEPA: Learning Interventional World Models with Graph Surgery Estimator
Abstract
Recent advances on Joint Predictive Embedding Architecture (JEPA) have shown significant progress in learning object-centric representations and the dynamics of these representations from images and videos. One prominent model is known as Causal-JEPA (C-JEPA) nam2026causaljepalearningworldmodels, which learns the interaction dynamics of objects by masking different feature maps based on images from previous time frames to predict the next time frame's representation state. The model is shown to be able to learn an important notion called influence neighborhoods, which are claimed to be the minimal predictor sets of each object in the image without causal sufficiency under mild assumptions. However, the relationship between influence neighborhoods and some classical causality concepts, such as invariant causal predictors, is not well-understood. In this work, we show that C-JEPA does not learn invariant causal predictors in a controlled environment in the presence of latent confounders. We propose an alternative training scheme to introduce an additional loss term that is based on graph surgery operations for learning object-level invariant predictors given a causal structure. Empirically, our proposed algorithm Do-CJEPA outperforms C-JEPA in learning the dynamics of the objects under mechanism shifts by having a smaller out-of-distribution test error gap and prediction errors in terms of mean squared error for both one-step prediction and rollout under shifted test environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.