Do-JEPA: From Masking to Intervention in Latent World Models
Abstract
Latent world models are trained to predict what happens next, so nothing in their objective separates what an action caused from what merely co-occurred with it. Object-masking models such as C-JEPA intervene on what the predictor can see; we intervene on what physically happens. From one saved simulator state we run the dynamics under an action and under a reference action , and train the model to predict the difference between the two latent futures. The resulting objective, Do-JEPA, has an effect loss, a support loss (where the action enters), a propagation loss (where its effect travels) and invariance losses (what must not change). In a synthetic system with object-aligned variables, support supervision finds the directly intervened object in of test cases, where a sparse action mask sends the action to a nuisance slot in every case, and response-onset supervision recovers the ring-shaped propagation graph (edge AUROC vs. ). From pixels, the effect loss beats a control trained on exactly the same data: it lowers latent effect error by on an end-to-end LeWM model ( seeds) and physical effect error by when trained and tested on natural action sequences, and on three independently generated CausalWorld benchmarks it lowers responsive effect error by about under physics shifts and the latent context sensitivity of predicted effects by . Trained from scratch it costs factual accuracy; fine-tuning an existing model with it removes this cost. Together, these results show that intervening on the world, rather than on what the model sees, helps latent world models predict what their actions cause.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.