acceptodds
Under review as a conference paper at ICLR 2027

3DFlowWAM: Bridging Video and Action through 3D Flow

Abstract

Robot actions are expressed in embodiment-specific coordinates, but their effects can be described through shared 3D geometry and motion. We introduce 3DFlowWAM, a world action model that uses 3D scene changes to connect video priors to robot control. Its central representation, 3D Flow, describes visible points through their 3D positions and displacements, preserving correspondence with video while remaining independent of robot action coordinates. By jointly predicting video, 3D Flow, and actions, the model learns to connect visual appearance with explicit spatial dynamics. We build on a pretrained video model through 3D Flow-aware mid-training, followed by task-specific post-training. We assess the contribution of 3D Flow through controlled comparisons on a common DreamZero backbone with matched video and action data. Experiments across simulation benchmarks and seven real-world tasks show that 3D Flow improves action learning both with and without mid-training, while mid-training further strengthens real-world adaptation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.