LaMo: A Latent Motion World Model for Long-Horizon Prediction
Abstract
Visual world models learned from video are often trained to predict future observations from past frames, but pixel-space forecasting is expensive and forces the model to allocate capacity to high-frequency details. Latent-space video models reduce this cost, yet they typically predict a dense latent state at every step, repeatedly carrying largely static information and making long-horizon rollouts harder to sustain. We propose , a probabilistic latent dynamics world model that predicts motion residuals—continuous latent variables encoding differences between consecutive motion tokens in a learned latent dynamics space. Given the current latent state and a short history of past motion tokens, our model autoregressively samples the next motion residual, integrates it with the previous motion token to obtain the next motion token, and recurrently advances the latent state using a frozen forward dynamics model. We instantiate the same contextual transformer backbone with two probabilistic parameterizations for residual prediction: Gaussian mixtures and conditional flow matching. We evaluate long-horizon rollouts on BDD100K using image-space metrics and semantic segmentation mIoU, showing improved long-horizon modeling in complex driving scenes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.