CL-WAM: Closing the Observation Loop in World–Action Models for Humanoid Manipulation
Abstract
World-action models (WAMs) predict video and generate actions, but refreshing an image context does not explicitly correct a persistent predicted state when contact outcomes diverge from a rollout. We introduce the Closed-Loop World–Action Model (CL-WAM), which closes this prediction–observation–action loop inside the world model. At each policy call, the previous belief, memory, and executed action produce a prior; the new observation provides an innovation that corrects the belief and updates memory before the next action chunk and, when scheduled, the next candidate rollout. A port-Hamiltonian-inspired latent field separates a skew/dissipative autonomous term from action-conditioned and predicted-contact terms, while scheduled rollout training feeds real and self-generated visual features through the same belief-update path. At deployment, slow candidate planning periodically refreshes a token that conditions a faster action head. Across four LIBERO suites, CL-WAM achieves 99.2% average success versus 98.6% for the officially released DiT4DiT checkpoint. On 24 RoboCasa-GR1 tasks, it reaches 67.50% (2430/3600) versus 56.7% (680/1200) for official DiT4DiT Run 1, a descriptive 10.8-point difference. Removing belief correction or scheduled rollout training lowers RoboCasa-GR1 success by 3.7 or 4.1 points, respectively. On seven real-world Unitree G1 tasks, observed success is 109/140 (77.9%) versus 101/140 (72.1%) for a locally matched DiT4DiT baseline.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.