acceptodds
Under review as a conference paper at ICLR 2027

MemJEPA: Erasure-Trained Memory Lets Latent World-Model Planners See Through Occlusion

Abstract

A latent world model can imagine the consequences of its actions in a learned state space, enabling model predictive control in robotics. However, an agent often loses sight of what it controls, such as when its arm moves out of the camera's frame. A windowed joint-embedding predictive architecture (JEPA) planner is then blind to the hidden actor, and adding memory trained only on visible clips does not help. We introduce Memory-Augmented JEPA (MemJEPA), which carries a recurrent belief through the planner's imagination and trains it with actor erasure: the actor is removed from the input clips but kept in the prediction targets. Theoretically, we show that a fixed-window latent planner cannot react to a hidden actor's motion that leaves no trace in its window, making memory of earlier frames necessary, and that erasure training against clean targets provides a learning signal for a useful carried state. Empirically, MemJEPA was tested on three image-based control tasks (reaching, pushing, navigation) with the frozen encoder of LeWorldModel (LeWM). MemJEPA raises average success under actor occlusion from 13% to  75% against a memoryless twin trained on the same data with the same actor erasure, while matching or exceeding its performance with the actor in view. The hidden-view gains transfer, at a smaller size, to a DINO-WM backbone.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.