Multi-Scale Future Modeling for Latent World Models
Abstract
Latent world models efficiently predict future world states in representation space, but most existing approaches model the future only at the timestep level, as a dense sequence of latent states. While this preserves fine-grained dynamics, it provides no explicit representation of future evolution over longer temporal intervals, potentially biasing forecasting toward local dynamics. We present MS-JEPA, a multi-scale latent forecasting framework that models future evolution at complementary temporal scales. The key idea is to represent the same future through both dense latent states and Future Abstractions: the former preserve fine-grained dynamics, while the latter summarize evolution over temporal intervals. A high-level predictor anticipates these abstractions from observed latents to guide dense forecasting. By explicitly modeling temporal structure across scales, our framework combines broader future context with fine-grained dynamics. Experiments across egocentric hand motion, whole-body motion, and physical-scene forecasting show consistent improvements over latent world modeling baselines, demonstrating the effectiveness and generality of our approach.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.