Learning Factorized World Transitions for Geometry-Aware Video World Models
Abstract
Video world models must preserve both physical structure and visual detail when predicting future observations. Pretrained geometry models provide valuable structural knowledge, but geometry alone does not specify the full visual prediction. This raises the question of how complementary representations should be organized for world-transition modeling. We introduce UPCAST, a framework that learns a geometry-anchored shared physical state together with a private appearance state. Cross-modal reconstruction relates transitions from different representations, while invertible coupling combines their shared and complementary information. A conditional teacher and selective knowledge transfer consolidate the learned representation into the original autoregressive video backbone, without auxiliary modules at inference. On camera-conditioned RealEstate10K, UPCAST reduces 64-frame FVD from 534.9 to 497.8 and JEDi from 4.40 to 3.88 relative to Geometry Forcing. Zero-shot ARKitScenes evaluation shows improved RGB fidelity with mixed sensor-referenced geometric outcomes under domain shift. An action-conditioned Minecraft evaluation further demonstrates applicability beyond camera control. Representation probes and training-stage analysis provide complementary evidence for the learned factorization and its transfer to standalone generation. Together, these results support organizing shared structure and complementary visual information within a predictive transition representation. Project page: https://upcast-anon.github.io/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.