HumanWorld: From Explorable to Human-centric Interactive World Model
Abstract
Recent advances in camera-controlled world models allow users to explore and navigate scenes and observe them from different viewpoints. However, viewpoint control alone does not enable users to act within a scene and change its state through interaction. We introduce HumanWorld, a human-centric interactive world model that enables users to explore a scene through camera control and actively change it by controlling the actions of a human character, including manipulating objects through actions such as picking them up and moving them. Our framework disentangles latent interaction dynamics from visual observation generation, with human motion and camera trajectories independently specified in a shared 3D world coordinate system. We learn a latent world interaction module to predict how controlled human motion changes object and scene states without directly synthesizing RGB observations. The interacted features then condition a pretrained video diffusion model through cross-attention to generate the corresponding visual observations. HumanWorld thus unifies viewpoint exploration, human motion control, and action-driven scene evolution within a framework for human-centric world modeling. Experiments across multiple datasets demonstrate that HumanWorld improves visual fidelity, human motion alignment, and human–scene interaction consistency over existing world models with the disentangled modeling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.