acceptodds
Under review as a conference paper at ICLR 2027

BODYWORLD: STREAMING HUMAN-CENTRIC WORLDS WITH INDEPENDENT BODY AND CAMERA CONTROL

Abstract

Human-centric video worlds require control over both what a person does and how the action is observed. We introduce BodyWorld, a framework that streams video from a reference image and independently editable body and camera commands. Decoupled geometric conditioning combines articulated body guidance with complementary scene appearance and viewing geometry. To maintain consistency as motion and viewpoint change, a static point-cloud memory anchors the scene, while Pose-Addressable History Memory recalls relevant generated observations to preserve human appearance beyond recent context. A three-stage curriculum learns joint control, introduces block-causal generation through pose-addressed teacher forcing, and distills few-step continuation with self forcing and distribution matching. Experiments across human motion, camera motion, pose recurrence, and long sequences show improved visual fidelity over prior methods and better appearance consistency across visibility changes. With cached history reuse, two-step inference reaches 13.36 generation FPS on a single NVIDIA H200 GPU.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.