An Action Is Worth One Patch: Unified World–Action Modeling with PatchWAM
Abstract
Robotic actions and their visual consequences describe the same motion: an action specifies how the robot moves, and a future frame shows where this motion leads, with the robot itself visible in the frame. We hypothesize that a predicted future frame substantially constrains the action, allowing the action to be read out of the predicted motion rather than generated by a second model. Most existing world action models nevertheless attach a separate action expert, which learns the same motion again from robot demonstrations alone. In this work, we introduce PatchWAM (Patch World-Action Model), which treats continuous actions as another type of patch through a fixed, parameter-free mapping called Action-as-Patch, so that the velocity field of a fine-tuned pretrained image generator produces the future frame and the action together. Under a linear probe, the future frame predicted by PatchWAM explains as much of the action’s motion as the ground-truth future frame, and clamping the future frame to a different frame during sampling changes the generated action accordingly. Under matched data and optimization settings, PatchWAM outperforms the dual-expert control at all five RoboTwin 2.0 data scales, with gains ranging from 5.8 to 14.3 percentage points. Under other training protocols, PatchWAM reaches 96.12% on RoboTwin 2.0 with additional augmented demonstrations and 91.8% on LIBERO-Plus.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.