EquiJEPA: Learning Action Representations for Effect Prediction and Action Generation
Abstract
An action representation should describe how an action changes the world and help an agent choose actions that produce a desired change. We introduce EquiJEPA, an inverse-dynamics Joint Embedding Predictive Architecture that learns this bidirectional connection in latent space. A state-blind effect head predicts latent changes from actions alone, encouraging transitions induced by the same action to concentrate around a shared prototype. Each transition consists of this prototype and a residual capturing state-dependent departures, including contact and collision. The prototypes provide concrete operations in latent space: an effect can be transferred to a new state, combined with another effect, and decoded into an action through the inverse head. Across simulated environments, we evaluate these operations through physical execution, including cross-state reuse and action combinations absent from training. EquiJEPA connects action-conditioned prediction with effect-conditioned action generation, providing an explicit action representation for reuse and compositional generalization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.