acceptodds
Under review as a conference paper at ICLR 2027

World4Act Policy: Learning Explicit 4D World Dynamics for Robot Policy Training

Abstract

World–Action Models improve robotic policy learning by coupling action prediction with forecasts of future visual dynamics. However, existing approaches predominantly supervise future evolution in RGB space, leaving the underlying 3D geometry and motion implicit. Although recent 4D world models incorporate geometric or motion cues, dense 3D structure and scene flow have yet to be systematically explored as action-conditioned prediction targets. Here we present World4Act Policy, an action-centred 4D World–Action Model that explicitly learns future visual appearance, 3D geometry and motion alongside robot actions. Given current multi-view RGB observations, robot states and language instructions, the model jointly predicts future action chunks, multi-view RGB frames, image-aligned point maps and 3D scene flow. To transfer pretrained video-generative priors to these 4D modalities, we develop dedicated geometry and motion variational autoencoders, together with a progressive training strategy that proceeds from 4D representation learning to world modelling and, ultimately, joint World–Action learning. For efficient deployment, we further introduce a modality-asymmetric causal architecture that decouples action generation from future visual and 4D tokens, such that all world-generation branches can be discarded at inference time. To provide the dense supervision required for learning action-conditioned 4D dynamics at scale, we also curate Robo4D_Dyn Dataset, a large-scale embodied-video dataset annotated with dense point maps and 3D scene flow. Experiments demonstrate that World4Act Policy predicts future 4D dynamics more accurately than representative 4D world models and achieves higher manipulation success rates than both video-based and 4D policy baselines, while preserving low-latency, action-only inference.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.