acceptodds
Under review as a conference paper at ICLR 2027

WARM: A World Action Rectification Model Unifying VLA and WAM through a Planning-Rectification Architecture

Abstract

Vision-Language-Action (VLA) models map instructions and images to actions and are strong at semantic planning, but relatively weak at contact-precise, physics-grounded control. World Action Models (WAMs) jointly generate future video and future actions, internalizing a physical forward model that can rectify motion, yet they struggle with long-horizon language planning. We show that the two families are complementary rather than competing: on 50 RoboTwin bimanual tasks they are only moderately correlated (Spearman ), each owns a disjoint set of signature successes, and the per-task oracle reaches , strictly above either family alone. We propose WARM (World Action Rectification Model), which unifies them through a plan-then-rectify architecture analogous to the cerebrum–cerebellum split: a frozen VLA plans a nominal trajectory , which is treated as an SDEdit prior and rectified by a flow-matching WAM. Ablations show that planning must come from the VLA (letting the WAM plan for itself drops success from to ), and that among coupling schemes (joint, joint_full, parallel, clamp), clamp—evaluate the plan while the video branch generates, then refine—is best. WARM attains on RoboTwin, improving over both the VLA () and WAM () baselines.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.