MetaDyn-JEPA: Learning to Predict across Physical Worlds with Interaction Evidence
Abstract
Physical interaction data often span systems with different underlying dynamics: the same action from a similar observed configuration can produce different outcomes. A short observation history may therefore be insufficient to predict the consequences of an action. We propose MetaDyn-JEPA (MetaDyn), a support-conditioned joint-embedding world model that uses additional interactions from the same physical system to inform prediction. MetaDyn encodes directed action–response transitions and aggregates them into a reusable world code that conditions multi-step latent prediction without test-time weight updates. Treating transitions as the units of evidence allows MetaDyn to aggregate support from disconnected fragments without relying on fragment-level sequence summaries. Across PokeWorld and MuJoCo, same-world support improves prediction over no-support baselines, while replacing it with wrong-world support degrades performance. MetaDyn remains stable under evidence-preserving repartitioning without fragment augmentation, and generally benefits from increasing same-world support toward the training length. In multi-dynamics Push-T, same-world support improves latent and physical prediction, and MetaDyn achieves higher reach-and-hold success than LeWM and LeWM + AdaJEPA. Overall, these results show that world-specific interaction evidence can provide a useful conditioning signal for prediction across varying dynamics.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.