DeltaWAM: Change-Centric Visual Foresight via Delta Tokens for Efficient World-Action Model
Abstract
World-Action Models (WAMs) offer promising foresight for robotic manipulation, yet efficient and dynamics-aware future representations haven’t been fully established. Pixel-space WAMs use representations that serve for repainting entire scenes, which brings up computational overhead, spatial-temporal redundancy and waste of model capacities on task-irrelevant information. To resolve these inefficiencies, we start from an overlooked observation: consecutive frames in physical manipulation usually share substantial contexts and differ only in structured and low-dimensional ways (e.g., object displacements and end-effector motion). In essence, it is these dynamic changes, rather than sequences of entire scenes, that an action policy needs to anticipate. Grounded in this principle, we build DeltaWAM, a World-Action model that shifts the fundamental unit of future predictions to a delta token. Each delta token is a single 1D vector that encodes the changes of dense DINO features between consecutive frames, inherently capturing dynamic transitions while reducing redundancy from shared static context. Our DeltaWAM builds on DeltaWorld, a latent world model that is pretrained on large-scale videos and forecasts change-centric futures: given a history of observed delta tokens, it autoregressively predicts the future delta sequence, one token per frame. These predicted transitions, alongside the DINO features of current observations acting as the spatial anchor, are then passed to a flow-matching action expert that generates the action chunk. Our experiments show that by focusing merely on what will change in the future, delta tokens prove to be more dynamics-aware and action-forcing, providing a natural bridge from high-level semantics to low-level control. With just 256 GPU hours on 2 H100 GPUs and 0.725B total parameters, DeltaWAM achieves high success rates(92.8%) on LIBERO while also demonstrating robust generalization under procedural perturbations on LIBERO-Pro. At inference the latency is 142.1 ms latency per action chunk and peak memory is only 3.86 GB.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.