Z-WAM: Learning World-Action Models from Image-Editing Priors
Abstract
Video-based world–action models (WAMs) connect robot action learning with dense future-video prediction, incurring the computational cost of synthesizing intermediate frames. We argue that this expensive visual prediction is unnecessary for learning effective actions: a policy can model the intended scene change and the actions that realize it without reconstructing every intermediate visual state. Motivated by this observation, we present Z-WAM, a vision–language–action framework that transfers pretrained image-editing priors to robot control. Conditioned on current observations and a language instruction, Z-WAM jointly predicts a single future multi-view observation and the corresponding action chunk. This formulation uses image editing to represent the outcome of a manipulation sequence while retaining the full sequence of executable actions. Sequential bidirectional attention couples the editing and action streams, allowing action hypotheses to refine visual features that subsequently guide action prediction. Progressive pretraining first adapts the editing model to robot observation transitions without action labels, then connects visual changes to controls through joint world–action learning. During task adaptation, grasp-aware objectives supervise cumulative arm motion and gripper transitions to encourage coordinated execution. Preliminary experiments on six real-world manipulation tasks support image-editing priors as a foundation for robot action learning without dense future-video prediction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.