Procedural Planning with JEPA World Model
Abstract
In this work, we introduce JEPA-Pro, a Joint-Embedding Predictive Architecture (JEPA)-based world model trained end-to-end for procedural planning, whose unit of prediction is a procedural step and in which text is used only to define actions and goals. World models predict how the state of the world changes under an action, which makes them a natural fit for model-based planning. Current world models, however, predict at a fixed temporal granularity which suits low-level control but is misaligned with the unit of decision in procedural tasks such as assembling furniture or cleaning a house. The challenge in procedural planning is twofold: (1) steps have variable duration ("wash the floor" may take seconds or minutes), and (2) they are rarely contiguous, since people often performs unrelated actions in between. To address this, our model is composed of two core components: an action-conditioned causal predictor trained jointly with the state encoder, and a learned cost function that estimates how far a predicted state is from the goal. At inference time, JEPA-Pro produces plans by test-time search over candidate action sequences in latent space, selecting the sequence whose predicted trajectory minimizes the cost, without generating pixels nor text. We evaluate JEPA-Pro on world modeling, procedural understanding and procedural planning benchmarks. We establish a new state-of-the-art on WorldPrediction-WM () and on EgoPlan-Bench (). On ENACT, the world model in zero-shot outperforms 23 of 30 VLM baseline, and all caption-free baselines on WorldPrediction-PP () using up to two order of magnitude less parameters. We show that finetuning the cost with few trajectories enables the world model to adapt to unseen tasks while naturally enabling test-time scaling during planning, outperforming zero-shot VLM policies on ALFRED and BEHAVIOR-1K.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.