acceptodds
Under review as a conference paper at ICLR 2027

No Gaussian Required: Contrastive Inverse Dynamics for JEPA World Models

Abstract

Joint-Embedding Predictive Architectures (JEPAs) learn world models by predicting future embeddings, but the objective admits a trivial solution of a constant encoder, so every practical system adds an anti-collapse mechanism. LeWorldModel (LeWM) prevents collapse with SIGReg, a regularizer that forces the latent distribution to match an isotropic Gaussian: the representation is stabilized by prescribing what it must look like, independently of the environment it models. We argue that the anti-collapse pressure can instead come from the transition data itself. Action-Contrastive Masked Transition Modeling (AC-MTM) keeps LeWM's forward latent prediction objective and adds a training-only inverse-dynamics head trained with Action-NCE: each latent transition must identify the action that produced it among the other actions in the batch, a discrimination task that a collapsed encoder provably fails. The inverse branch is discarded after training, leaving test-time encoding, forward prediction, planning, and compute identical to LeWM. On four standard pixel-control tasks under a matched planning protocol, AC-MTM trains stably from scratch and matches or exceeds SIGReg. On the harder multi-object OGBench Visual Scene task, results suggest that SIGReg's prescribed geometry is a bottleneck: AC-MTM reaches success rate versus for SIGReg, improving by 20–24 points in each training seed. A random-policy run gives a 52% baseline estimate. Contrastive inverse dynamics thus provides a distribution-free anti-collapse signal that requires no target network, stop-gradient, pretrained encoder, or reconstruction objective, and we characterize the action space and observability assumptions under which it holds. We will release code and configuration files after review.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.