ScaleWorld: Streaming High-Resolution World Generation through Shared Prediction and Refinement
Abstract
High-resolution generation is crucial for camera-controlled world models, yet direct high-resolution rollout is costly and often degrades visual details. We present , a causal video model that learns world prediction and high-resolution refinement in shared parameters. This allows the model to refine its own predictions for improved visual detail while reducing the cost of high-resolution generation. Specifically, each low-resolution chunk is completed through four-step camera-controlled prediction and then conditions a second invocation of the same backbone for one-step refinement. Realizing this process requires addressing two core challenges: (i) Efficiency. Separate causal states maintain a long prediction context at low resolution and a short refinement context at high resolution, reducing context cost. (ii)Capability integration. We separate specialist optimization from consolidation and introduce task-asymmetric on-policy distillation (TA-OPD), which transfers prediction and refinement behaviors into a shared student under their respective sampling and history policies. Extensive experiments show that effectively unifies prediction and refinement while reducing the context-storage cost of high-resolution world generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.