Fast LeWorldModel
Abstract
Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models. For visual planning, however, LeWM evaluates candidate action sequences by repeatedly applying a local one-step latent transition model. This autoregressive rollout makes planning computationally expensive and causes accumulated latent prediction errors as the horizon grows. We propose We propose Fast LeWorld-Model (Fast-LeWM), a fast latent world model that replaces repeated local rollout with action-prefix prediction. Given the current latent state and a candidate action sequence, Fast-LeWM encodes its prefixes and predicts the future latents reached after executing those prefixes in parallel. By making action prefixes the basic prediction unit, Fast-LeWM directly models action effects accumulated over different horizons. This prefix-level supervision forces Training these multi-horizon predictions jointly encourages the model to learn how states continuously evolve under different action prefixes, rather than only fitting one-step state transitions. During planning, the predictor uses the prefix from the encoded action sequence to evaluate the corresponding future latent directly, without explicitly rolling through each intermediate imagined state. Across multiple tasks, Fast-LeWM improves average success over LeWM while substantially reducing planning time. It also achieves lower open-loop latent prediction loss, with errors growing considerably more slowly as the rollout horizon increases.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.