SWAN: Scaling Long-horizon World Models via Next-Scale Autoregression
Abstract
Robotic world models must faithfully predict the consequences of prescribed actions, including failures, while preserving physical states over extended horizons. However, existing generative models frequently hallucinate success due to strong visual priors, and struggle with error accumulation or prohibitive memory growth during recursive rollouts. To overcome these challenges, we introduce SWAN, which Scales World modeling through Autoregressive Next-scale prediction by integrating unified action-conditioned prediction with a coarse-to-fine hierarchical streaming memory. First, scale-aligned action representations guide generation across all resolutions to strengthen action-future coupling. Second, a time-scale memory preserves fine recent and coarse distant causal states under a bounded resident-memory budget, coupled with scale-wise dream forcing to bridge the exposure gap. Finally, we align visual dynamics on fresh rollouts via action-following and consistency rewards under fixed controls, without task-success rewards. Experiments demonstrate stable streaming rollouts up to 60 seconds. Across 600 matched sim-to-real rollout pairs with frozen external policies, SWAN achieves a success-rate MAE of 0.062 and a within-task rank correlation of 0.88, while an auxiliary action expert further enables effective real-robot control. Despite residual optimism in contact-rich stages, these results demonstrate the promise of next-scale autoregression for scalable robotic simulation and execution.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.