-WM: Predicting Steps at Once for Long-Horizon Planning with World Models
Abstract
Latent world models have become an appealing approach to visual control, enabling goal-conditioned optimization at test time and generalization across tasks with a single model. In practice, however, planning success remains limited to short horizons, and gradient-based planners driven by autoregressive rollouts still underperform sampling-based methods, despite their potential to optimize action sequences efficiently through model gradients. We hypothesize that two properties of next-step prediction contribute to these limitations: autoregressive rollouts compound prediction errors, and planning gradients must traverse a chain of nested model calls. We study a simple alternative: predicting the entire future sequence in a single forward pass. We introduce -WM, a minimal extension to transformer-based JEPA world models: a multi-horizon training objective and input layout that let the predictor map a variable-length action sequence directly to the corresponding sequence of future latent states in a single forward pass. We apply -WM to three classes of JEPA world models and evaluate on four environments. Overall, -WM improves gradient-based planning by 37% and sampling-based planning by 15% over the existing next-step world models it extends. Our analysis shows that the gains cannot be explained by parallel rollout alone: improvements persist under autoregressive inference and coincide with substantial changes in the learned representation. Parallel prediction additionally reduces long-horizon inference cost. Together, -WM offers a simple, broadly applicable path to more accurate and efficient long-horizon control. Videos are available at https://anon-k-wm.github.io.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.