acceptodds
Under review as a conference paper at ICLR 2027

OpenWorld: Scaling World Model Pretraining

Abstract

World models trained on large-scale data hold immense promise for emulating complex interactions with the physical world. While action-conditioned video generation as a world model has gained traction due to its potential to scale, the underlying scaling behaviors of these models remain poorly understood. Most existing works rely on finetuning fixed-size pretrained text-to-video models on action data; however, a systematic understanding of how model size, data, and compute affect action-conditioned video world models is lacking. To close this gap, we curate an action-video dataset comprising 4 million video clips (120 billion image tokens) spanning 10,000 hours of human, robot, and navigation data. We then introduce a fully open recipe, OpenWorld, that enables scaling models from 600M to 7.5B parameters. In a dedicated isoFLOP sweep matched to our deployment-scale training setup, we find that held-out diffusion loss exhibits a strongly model-heavy compute frontier, with . Separately, a constant-learning-rate training-loss frontier fitted on models up to 1.02B parameters predicts the excluded 7.5B run to within 2.5% near its predicted optimal compute. Rollout evaluations show that larger deployment-scale checkpoints improve in-domain video fidelity, downstream finetuning, and behavior-level trajectory matching, while also revealing that standard paired video metrics are not always reliable proxies for action following. Finally, we demonstrate the utility of OpenWorld across downstream applications, including real-time human-interaction simulation, multi-view robot execution simulation across diverse robot morphologies, and building a robot leaderboard for evaluating policies in the cloud. Qualitative rollouts from OpenWorld-7.5B further show plausible gravity, contact-rich interaction, and deformable-object dynamics.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.