Generative Cinematographer: Composing Camera and Object Motion in 3D
Abstract
Existing controllable video generation methods can model camera motion or object motion, but most operate through image-space controls such as 2D trajectories or sparse drag signals. These controls are fundamentally ambiguous: the same image-plane motion may correspond to multiple underlying 3D motions, making camera motion and object motion difficult to disentangle under viewpoint changes. We present Generative Cinematographer (GenCine), a controllable video generation framework that lifts a single image into an editable 3D scene scaffold where camera motion and foreground motion are authored jointly in a shared world frame. Our key observation is that world-space motion can be rasterized into dense correspondence maps that preserve persistent identities across viewpoints while remaining compatible with RGB-pretrained video diffusion models. The resulting representation supports sparse local 3D motion handles, coordinated camera–object choreography, and piecewise-rigid approximations of non-rigid motion without requiring physics simulators or category-specific priors. To preserve the generative prior of a pretrained video model, we freeze a Wan-based diffusion backbone and condition it through lightweight adaptation layers driven by the proposed correspondence representation. Experiments demonstrate coherent camera-relative motion, improved geometric consistency under viewpoint changes, and strong controllability across diverse real-world scenes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.