ASTER: Universal World Simulation via Scalable Interaction Pretraining
Abstract
Interaction data offer a way to train video world simulators to predict the consequences of actions, but action representations differ across robots, human videos, and simulators. We present ASTER, a video world simulator conditioned on an initial frame, point flow on a moving body, and camera geometry. Only the mover's motion is specified; the surrounding scene response must be predicted. This shared input lets us pretrain an 8B video diffusion transformer on 13 real and simulated interaction sources. Across eight evaluation settings, ASTER reduces perceptual prediction error by 34% on average relative to the strongest external predictor and by 16% relative to its backbone, Cosmos 3. After adaptation to painting and laser cutting, it reaches the requested target in 83% of cases and produces a valid visible process in 73%, compared with 42% and 17% for a text-conditioned image-to-video baseline. At fixed model size, prediction improves across a 230× pretraining compute range, and the full robot, human, and object mixture lowers perceptual error relative to robot-only pretraining in all six evaluated source domains.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.