World-4D: 4D Embodied World Models for Robotic Manipulation
Abstract
Robotic manipulation in the open world requires not only recognizing what a scene looks like, but also anticipating how its 3D structure moves under interaction. We argue that synchronized RGB, depth, and optical flow (RGB-DF) provide a geometry-aware projective representation that makes scene structure and pixel-level motion explicit. Compared to 2D pixel videos, this multi-modal synergy aligns visual appearance with geometric structure and temporal motion, creating a representation space significantly closer to low-level end-effector actions demanded by robotic systems, narrowing the gap between world prediction and policy learning. Building on this insight, we introduce World-4D, a generative model that co-produces future RGB frames, depth maps, and optical flow from a single RGB-D image and a language instruction within one unified diffusion process. This 4D world model features a tri-branch architecture that integrates cross-modal attention with frame-wise 3D RoPE, ensuring that appearance, geometry, and motion evolve consistently. To supply training data at scale, we curate World-4D-200M, a large-scale dataset containing approximately 254 million frames across egocentric human and robotic manipulation videos with pseudo-labels for depth and optical flow. We further propose World-4D-Policy, an inverse dynamics head that consumes the internal 4D representations of World-4D in a single forward pass, bypassing expensive multi-step denoising, to output robot actions in a closed-loop manner. Experiments show that World-4D produces temporally and spatially coherent 4D predictions, and that World-4D-Policy achieves strong performance on real-world bimanual manipulation tasks, outperforming the evaluated baselines on most tasks, with notable gains on tasks demanding spatial precision and temporal coordination.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.