CausalWAM: Learning Action-Conditioned Visual Dynamics for Efficient World Action Models
Abstract
World Action Models (WAMs) jointly model robot actions and future visual observations, leveraging future visual dynamics as dense supervisory signals to learn physically grounded action and alleviate the sparse action-supervision bottleneck in Vision-Language-Action (VLA) models. However, this paradigm suffers from : although WAMs may generate plausible future observations, their predictions remain dominated by the current visual context and learned video priors, exhibiting only weak dependence to the conditioning action. To address this issue, we present , a causally grounded Mixture-of-Transformers (MoT) framework with a mixed training recipe. The recipe combines Action-Conditioned World Modeling (AC-WM), which uses clean action tokens to provide direct supervision for visual dynamics, with World Action Modeling (WAM) training, which preserves the joint action–video modeling pathway required for action generation. A progressive schedule is developed to emphasize AC-WM early to establish reliable action dependence and gradually shifts toward WAM to integrate the learned dynamics into action generation. By treating actions as conditioning variables rather than imitation targets, CausalWAM effectively exploits failed and suboptimal trajectories as transition-level supervision, thereby enhancing data efficiency. Experiments in simulation and on real robots demonstrate the effectiveness of CausalWAM in improving action dependence and policy performance, while system-optimized action-only inference reaches on an NVIDIA RTX . The code will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.