Who Acts, Where, and How: Unified World Model With Heterogeneous Controls
Abstract
We introduce a unified video world model for compositional control over viewpoint, embodiment, and action. Joint learning is hindered by inconsistent control representations and sparse paired supervision. We align partially annotated controls from 18 datasets spanning robot demonstrations, human egocentric recordings, and camera-motion videos through reference-camera geometry, a unified action template, and a shared action encoder. This provides consistent physical semantics for learning from complementary sources. Visual context supplies embodiment cues without requiring explicit identifiers at inference. A hybrid Mixture-of-Transformers and Mixture-of-Experts architecture couples pretrained vision-language conditioning with noise-specialized generation. We pretrain with supervised image-to-video learning and cross-segment self-supervised video prediction, then adapt the high-noise expert using limited cross-view pairs for video re-shooting. The low-noise expert remains frozen, preserving refinement learned from heterogeneous pretraining. Experiments show strong visual fidelity and control accuracy in human and robot action-conditioned generation, camera-controlled generation, and video re-shooting. Qualitative examples further illustrate compositional action–camera control and human-to-robot transfer.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.