acceptodds
Under review as a conference paper at ICLR 2027

AGR-JEPA: Action-Grounded Residual Joint-Embedding Predictive Architecture for Planning

Abstract

In this research, we propose Action-Grounded Residual Joint-Embedding Predictive Architecture (AGR-JEPA), a reconstruction-free latent world model for planning. Our work is based on a key observation: existing approaches typically trained to minimize next-embedding prediction error, whereas effective planning additionally requires the consequences of different actions to remain distinguishable in latent space. Through controlled experiments, we find that lower prediction error does not consistently correspond to higher planning success, indicating that latent predictability alone is not a reliable proxy for planning utility. We refer to this mismatch as the prediction–planning gap. To address this gap, AGR-JEPA adopts a residual parameterization that represents the future latent state as the combination of the current latent state and an action-conditioned latent update, focusing dynamics learning on action-induced state changes. Furthermore, AGR-JEPA introduces a control-grounding objective that uses action recoverability as a training signal, encouraging latent updates to retain information relevant to their conditioning actions and thereby improving the behavioral discriminability of action outcomes. This auxiliary objective is used only during training and incurs no additional computation at planning time. Across multiple 2D and 3D control tasks, AGR-JEPA achieves an average planning success rate of 92.06%, delivering consistent improvements over leading baselines across multiple tasks.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.