M-WAM: Motion-Centric World Action Model
Abstract
World-action models learn robot control alongside future prediction, but their predictive objectives commonly emphasize visual appearance. We investigate dense 3D scene motion as a supervision target for action learning. We introduce M-WAM, which predicts pixel-aligned 3D point trajectories organized into image-like motion maps. A pretrained image-generation backbone models these trajectories, while motion and action objectives jointly optimize the scene features used by the policy. In training, we reconstruct dense trajectories on demand from the per-frame simulator geometry. This avoids precomputing every frame pair and reduces export and storage costs from quadratic to linear in episode length. At deployment, the policy generates actions directly without predicting future motion. In RoboTwin ablations, 3D motion improves success by 10.9 percentage points over last-frame RGB prediction and 6.9 points over 2D motion. M-WAM's mean success rate on LIBERO-Plus is 1.7 points higher than ImageWAM's. Across three real-robot tasks, it achieves twice the mean success rate of LingBot-VA. These results support explicit 3D motion prediction as an effective training objective for manipulation policies. All code and data will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.