WARM: A World Action Rectification Model Unifying VLA and WAM through a Planning-Rectification Architecture
Abstract
Vision-Language-Action (VLA) models map instructions and images to actions and are strong at semantic planning, but relatively weak at contact-precise, physics-grounded control. World Action Models (WAMs) jointly generate future video and future actions, internalizing a physical forward model that can rectify motion, yet they struggle with long-horizon language planning. We show that the two families are complementary rather than competing: on 50 RoboTwin bimanual tasks they are only moderately correlated (Spearman ), each owns a disjoint set of signature successes, and the per-task oracle reaches , strictly above either family alone. We propose WARM (World Action Rectification Model), which unifies them through a plan-then-rectify architecture analogous to the cerebrum–cerebellum split: a frozen VLA plans a nominal trajectory , which is treated as an SDEdit prior and rectified by a flow-matching WAM. Ablations show that planning must come from the VLA (letting the WAM plan for itself drops success from to ), and that among coupling schemes (joint, joint_full, parallel, clamp), clamp—evaluate the plan while the video branch generates, then refine—is best. WARM attains on RoboTwin, improving over both the VLA () and WAM () baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.