acceptodds
Under review as a conference paper at ICLR 2027

WorldFlow: Joint Video and 3D Motion Prediction via Large-Scale Pretraining

Abstract

Predicting future physical states stands as a central challenge in embodied intelligence. Existing video world models leave 3D dynamics implicit, action-based formulations remain tied to specific robot configurations, and 3D motion prediction models use flow-only supervision, providing a limited learning signal for physical dynamics. To circumvent these limitations, we introduce WorldFlow, a generative world model that jointly predicts future video and query-based 3D trajectories. By parameterizing 3D motion as continuous displacement tokens aligned with the video grid, a unified Mixture-of-Transformers (MoT) backbone jointly denoises both modalities under a rectified-flow objective. This formulation couples explicit geometric dynamics with visual evolution, effectively unlocking the physical commonsense embedded within pretrained video models. To empirically scale and validate this unified architecture, we train WorldFlow on a heterogeneous mixture of 1.44 million object-level samples spanning simulated and real-world domains. We further introduce WorldFlowBench, a benchmark comprising three evaluation sets, including SpatialCog for assessing spatial reasoning. Across all three evaluation sets, WorldFlow achieves the lowest ADE and FDE among the evaluated learning-based methods. Crucially, on SpatialCog, our joint model reduces Average Displacement Error (ADE) by 55.9% and Final Displacement Error (FDE) by 68.5% relative to a flow-only ablation, demonstrating that generative visual priors fundamentally regularize complex geometric forecasting. Our code, pretrained models, and benchmarks will be made publicly available.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.