ViSPaC: Visual Spatial Pose Codec for Camera-Controlled Video Generation
Abstract
Camera encodings and scene geometry have enabled substantial progress in camera-controlled video generation. We view camera control as a representation problem of relating geometric motion to visual dynamics, and propose ViSPaC, a Visual Spatial Pose Codec for joint modeling of camera motion and video. ViSPaC represents camera trajectories as evolving projections of colored landmarks in a synthetic 3D reference structure. These projections expose the visual effects of camera motion, while landmark identities and image positions establish correspondences for geometric recovery, without requiring reconstruction of the underlying scene. We identify two coupled constraints on this representation: visual compression limits the number of reliably distinguishable landmark identities, while their spatial distribution determines coverage across camera views. ViSPaC separates color identity from spatial placement and uses a spatial codebook to adapt the reference structure's scale and layout within a limited identity budget. A shared video diffusion transformer models video and visual pose, supporting both camera-controlled generation and video-conditioned pose recovery. Controlled comparisons show that ViSPaC improves camera control and pose recovery over alternative representations, with further gains from its identity budget and spatial adaptation. On DL3DV and Tanks and Temples, ViSPaC outperforms both implicit camera-conditioning and explicit scene-geometry methods, reducing relative translation and rotation errors by an average of 24.7% and 31.7%, respectively, over the previous state of the art.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.