acceptodds
Under review as a conference paper at ICLR 2027

Preserving Unstable Modes Through Inverse Dynamics in JEPA World Models

Abstract

Robotic systems often exhibit unstable modes, along which small perturbations and disturbances can cause unbounded growth unless corrected through feedback. Controlling such systems from high-dimensional visual observations requires representations that preserve these modes. Joint-embedding predictive architectures (JEPAs) provide a natural framework for learning representations and their dynamics from visual data. However, we demonstrate that next step prediction combined with anti-collapse regularization provides no such guarantee: the training loss can be minimized while controllable unstable modes are collapsed, making stabilization from the learned representation impossible. To circumvent this, we augment predictive world-model training with an action reconstruction objective (i.e., an inverse dynamics loss) that encourages control-aware representations, namely, visual representations that preserve crucial features for control. We prove that exact action reconstruction makes the encoder injective on the finite-horizon reachable subspace, hence, no state direction produced by an action sequence within steps can be discarded. Moreover, as grows, the dominant eigenspace of the finite-horizon controllability Gramian reveals the controllable unstable subspace. Despite our theory considering linear systems, experiments on nonlinear visual control tasks (CartPole, Walker2D, and PointMaze) validate our findings and show the benefits of control-aware representation learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.