acceptodds
Under review as a conference paper at ICLR 2027

An Action Is Worth One Patch: Unified World–Action Modeling with PatchWAM

Abstract

Robotic actions and their visual consequences describe the same motion: an action specifies how the robot moves, and a future frame shows where this motion leads, with the robot itself visible in the frame. We hypothesize that a predicted future frame substantially constrains the action, allowing the action to be read out of the predicted motion rather than generated by a second model. Most existing world action models nevertheless attach a separate action expert, which learns the same motion again from robot demonstrations alone. In this work, we introduce PatchWAM (Patch World-Action Model), which treats continuous actions as another type of patch through a fixed, parameter-free mapping called Action-as-Patch, so that the velocity field of a fine-tuned pretrained image generator produces the future frame and the action together. Under a linear probe, the future frame predicted by PatchWAM explains as much of the action’s motion as the ground-truth future frame, and clamping the future frame to a different frame during sampling changes the generated action accordingly. Under matched data and optimization settings, PatchWAM outperforms the dual-expert control at all five RoboTwin 2.0 data scales, with gains ranging from 5.8 to 14.3 percentage points. Under other training protocols, PatchWAM reaches 96.12% on RoboTwin 2.0 with additional augmented demonstrations and 91.8% on LIBERO-Plus.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.