SyncWorld: Many Views, One World
Abstract
An interactive world observed by multiple players must support different views of the same evolving scene. We introduce a view-decoupled architecture that separates shared world evolution from observer-specific video generation. Joint actions advance an authoritative state, geometric projections express that state in each observer's image plane, and one shared video model renders the observations. Adding an observer reuses the learned weights through another renderer call. We instantiate the architecture with engine-authoritative dynamics and a video-diffusion backbone adapted through geometric and status conditioning, gated residual adapters, and LoRA. We also construct a CS:GO dataset pairing multi-observer video with cameras, actions, events, and state. We will release the five-map raw collection and 84,066 processed Dust2 clips, approximately 118 cumulative hours. To evaluate cross-view geometric compatibility, we introduce SymMVC, a symmetric reprojection residual using dense correspondences and calibrated cameras. Our renderer achieves 24.53 dB PSNR versus 15.53 dB for additive residual conditioning under matched training data and budget, and a 480-pixel SymMVC residual of 0.5259 versus 5.1 under shuffled state. A native trajectory drives ten concurrent views, and autoregressive evaluation reaches 12 windows, approximately 60 seconds.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.