acceptodds
Under review as a conference paper at ICLR 2027

ManiJEPA: Manifold-Aware JEPA World Models via Directional Contraction

Abstract

Joint-embedding predictive architectures learn world models by predicting future representations from current ones, enabling planning directly from pixels. But when the predictor is rolled out autoregressively at planning time, small errors compound exponentially, making long-horizon control unreliable. Existing fixes either leave the exponential form unchanged or force the predictor's Jacobian to uniformly contract, bounding error at the cost of destroying the sensitivity to early actions that planning requires. We show that this tension has a principled resolution: constrain only the directions that the training data cannot identify. Because encoder latents from reachable observations lie on a low-dimensional manifold, the predictor's behavior normal to the manifold is never observed by the teacher-forced loss and can be regularized at no bias cost, while the tangent directions that carry dynamics and control signal are preserved. ManiJEPA estimates the local tangent space of the latent manifold via local PCA on neighboring latents, decomposes the predictor Jacobian into tangent and normal blocks, and imposes three hinge penalties: normal contraction, a tangent shell that preserves on-manifold gain, and a leak term that blocks error from flowing back into the tangent space. An explicit feasibility condition coupling these penalties guarantees at most linear rather than exponential error growth, and a control-authority lower bound ensures early actions retain influence over the plan. On PushT, TwoRoom, OGBench-Cube, and LIBERO-Goal, ManiJEPA matches baselines at short horizons and substantially outperforms them at long horizons, with ablations confirming that each directional component is necessary.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.