CanoCam: Gimbal and Intrinsics Control for Video Generation via Canonical Views
Abstract
Camera-controlled video generation remains unreliable under extreme pitch and roll, wide fields of view, and strong lens distortion. These camera configurations are rare in training data, making their visual effects difficult to learn jointly with scene dynamics. We present CanoCam, which reduces this learning burden through canonical video generation and geometric reprojection. Our approach exploits a geometric distinction: translation induces depth-dependent parallax and visibility changes, whereas rotation and changes in intrinsics admit a depth-independent ray mapping at a fixed camera center. CanoCam first generates canonical videos under simplified camera conditions while preserving the requested camera-center trajectory. We consider two canonical representations: level perspective views that retain yaw and panoramic views that provide full angular coverage. Reprojection then applies the remaining rotation and target intrinsics. For both representations, a target-camera-conditioned video model completes missing regions and refines the reprojected content. Experiments on SpatialVID-Extreme, Prompt Entanglement, and PanShot show improved camera control over strong baselines across extreme orientations and diverse camera models, including both pinhole and fisheye cameras. Controlled ablations validate canonicalization and geometric reprojection as a simple and effective approach to robust camera control. The code will be made publicly available upon acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.