PA-WAM: Progress-Aligned World–Action Model for Long-Horizon Robot Manipulation
Abstract
World–Action Models (WAMs) provide dynamics supervision for robot policies by predicting future states, but in long-horizon manipulation, a small future-representation error does not necessarily imply correct task-stage advancement.Existing future-reconstruction objectives implicitly contain stage changes in demon-strations, yet do not explicitly distinguish how different prediction biases affect task progress. We propose Progress-Aligned World–Action Model (PA-WAM), which introduces history-conditioned stage-transition supervision into LaWAM’s latent future learning. PA-WAM first aligns successful demonstrations for the same task to construct a continuous stage coordinate, and uses current manipulation evidence and multi-scale history to predict the nominal transition within a fixed horizon;then, Bidirectional Progress–Dynamics Coupling (BPDC) uses this transition to modulate the latent action query and constrains the stage change of the generated future through a frozen progress readout. The target–readout consistency objective is jointly optimized with future reconstruction and action imitation, while the readback branch is used only during training. Compared with LaWAM, PA-WAM improves the average success rate on LIBERO from 98.6% to 99.2%, with the largest gain on Long, from 97.0% to 98.2%; on MimicGen-D1, it improves from 81.8% to 84.3%. Across three RM65 real-robot tasks, the average success rate increases from 82.2% to 87.8%. These results indicate that treating task progress as an internal constraint on future generation can improve the execution capability of world–action models in long-horizon robot manipulation while preserving local dynamics modeling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.