ShiftWM: World Models that Move What They Saw
Abstract
World models let an agent evaluate actions before executing them, and forecasting in the patch-feature space of a frozen visual foundation model keeps them compact. Current latent world models, such as DINO-WM and V-JEPA 2-AC, regenerate the whole feature grid at every future step and feed their predictions back as inputs. Yet in manipulation and surgical scenes most content persists and only moves, so regeneration corrupts the static background and lets errors compound. We introduce ShiftWM, a latent world model that forecasts by moving what it has already seen: for every future step it predicts an action-conditioned soft transport of nearby observed patch features, learned without flow supervision, a gate between moved and kept features and a correction for new content, and it decodes all horizons in parallel. As an output head, it also attaches to open-source world models. On held-out real-robot video from DROID, Language-Table and the bimanual IWS tasks, and on seven surgical tasks from Open-H, ShiftWM has the lowest error of all matched output heads, 7.6% below a direct predictor on DROID; as a plug-in, it lowers the error of the 1.3B-parameter V-JEPA 2-AC by 15.7%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.