DualWAM: Dual-System World Action Models for Asynchronous Global Planning and Local Refinement
Abstract
World Action Models (WAMs) jointly generate robot actions and predict future world states, transferring priors from video pretraining to robot control. However, future visual prediction is computationally expensive, so existing WAMs often rely on long action chunks to amortize inference cost across control steps, at the cost of closed-loop responsiveness. We present DualWAM, a dual-system WAM that preserves broader-horizon world-action generation while enabling high-frequency closed-loop action updates by decoupling global planning and local refinement. System 2 periodically performs high-noise bidirectional denoising over a broader world-action chunk to establish a global plan, while wrist-only System 1 extracts a temporally aligned short window from the intermediate denoising state and completes low-noise refinement using the latest wrist observations, which provide action-aligned cues about local geometry, motion, and contact during interaction. The two systems operate asynchronously along a shared denoising trajectory: each global plan is reused across multiple local updates, while System 1 repeatedly incorporates fresh interaction feedback. Across zero-shot manipulation tasks on Franka and Galbot, DualWAM improves success over the strongest evaluated baseline by 4.5 percentage points on average, while achieving a 16.6 critical-path speedup. Further studies show that role-matched egocentric and UMI data improve success by 14 percentage points, and that the decoupled design naturally supports edge–cloud deployment with substantially lower communication overhead than the baseline.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.