3DFlowWAM: Bridging Video and Action through 3D Flow
Abstract
Robot actions are expressed in embodiment-specific coordinates, but their effects can be described through shared 3D geometry and motion. We introduce 3DFlowWAM, a world action model that uses 3D scene changes to connect video priors to robot control. Its central representation, 3D Flow, describes visible points through their 3D positions and displacements, preserving correspondence with video while remaining independent of robot action coordinates. By jointly predicting video, 3D Flow, and actions, the model learns to connect visual appearance with explicit spatial dynamics. We build on a pretrained video model through 3D Flow-aware mid-training, followed by task-specific post-training. We assess the contribution of 3D Flow through controlled comparisons on a common DreamZero backbone with matched video and action data. Experiments across simulation benchmarks and seven real-world tasks show that 3D Flow improves action learning both with and without mid-training, while mid-training further strengthens real-world adaptation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.