acceptodds
Under review as a conference paper at ICLR 2027

ScaleWorld: Streaming High-Resolution World Generation through Shared Prediction and Refinement

Abstract

High-resolution generation is crucial for camera-controlled world models, yet direct high-resolution rollout is costly and often degrades visual details. We present , a causal video model that learns world prediction and high-resolution refinement in shared parameters. This allows the model to refine its own predictions for improved visual detail while reducing the cost of high-resolution generation. Specifically, each low-resolution chunk is completed through four-step camera-controlled prediction and then conditions a second invocation of the same backbone for one-step refinement. Realizing this process requires addressing two core challenges: (i) Efficiency. Separate causal states maintain a long prediction context at low resolution and a short refinement context at high resolution, reducing context cost. (ii)Capability integration. We separate specialist optimization from consolidation and introduce task-asymmetric on-policy distillation (TA-OPD), which transfers prediction and refinement behaviors into a shared student under their respective sampling and history policies. Extensive experiments show that effectively unifies prediction and refinement while reducing the context-storage cost of high-resolution world generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.