acceptodds
Under review as a conference paper at ICLR 2027

ASTER: Universal World Simulation via Scalable Interaction Pretraining

Abstract

Interaction data offer a way to train video world simulators to predict the consequences of actions, but action representations differ across robots, human videos, and simulators. We present ASTER, a video world simulator conditioned on an initial frame, point flow on a moving body, and camera geometry. Only the mover's motion is specified; the surrounding scene response must be predicted. This shared input lets us pretrain an 8B video diffusion transformer on 13 real and simulated interaction sources. Across eight evaluation settings, ASTER reduces perceptual prediction error by 34% on average relative to the strongest external predictor and by 16% relative to its backbone, Cosmos 3. After adaptation to painting and laser cutting, it reaches the requested target in 83% of cases and produces a valid visible process in 73%, compared with 42% and 17% for a text-conditioned image-to-video baseline. At fixed model size, prediction improves across a 230× pretraining compute range, and the full robot, human, and object mixture lowers perceptual error relative to robot-only pretraining in all six evaluated source domains.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.