JAM: Effective Coupling of World and Action in World-Action Model Pretraining
Abstract
World Action Models (WAMs), which jointly predict future observations and robot actions, have emerged as a promising paradigm for generalist robot control. Their effectiveness is closely tied to how world and action predictions are coupled during pretraining. Existing WAMs condition one modality on the other, merge both into a shared token sequence, or use multi-stream joint attention between world, action, and text. However, these designs restrict information to one direction, let the high-dimensional world signal dominate the compact action signal, or entangle direct world–action coupling with semantic understanding. We introduce JAM (Joint World–Action Modeling), a WAM pretraining framework where world and action streams exchange information and gradients through MM-DiT-style joint attention, while text conditions both streams. JAM further uses 2D motion tracks to provide an embodiment-agnostic image-space target that makes world–action correspondence explicit and extends supervision to action-free human videos. Pretrained primarily on action-labeled robot data, JAM scales consistently with compute and model capacity. Evaluations across seven simulation benchmark families and real-world experiments demonstrate that the same pretrained JAM model adapts to single-arm, bimanual, mobile, and humanoid embodiments and supports action-only inference without generating future observations. Controlled ablations further show that our two-stream world–action coupling outperforms one-way cross-attention and a shared-residual joint-attention baseline, and that motion-track supervision improves downstream performance.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.