Camera-Controlled Video Generation at Scale
Abstract
Camera-controlled video generation has become a promising interface for cinematic content creation, but scaling up 3D-camera-pose-conditioned training remains challenging. In large video corpora, camera translation is typically not aligned to a metric scale, making global motion control ambiguous across examples. Metric-alignment approaches address this issue with a preprocessing pipeline that combines structure-from-motion and metric depth estimation to approximate a shared scale. Applying this pipeline to every training video makes metric alignment a computational bottleneck as the corpus grows, while depth-estimation errors can also affect the aligned scale. To remove this bottleneck, we propose a scale-decoupled formulation for camera conditioning that treats trajectory magnitude as an explicit control variable rather than a quantity that must be metrically calibrated in advance. We normalize translation by its total path length, retaining zero translation for stationary cameras. This preserves direction and relative frame-to-frame magnitude while discarding unreliable global scale. Our goal is scalable training without metric-scale alignment, rather than resolving monocular scale ambiguity. We then inject camera-pose and scale , where controls the overall strength of camera translation, denotes a sinusoidal-style embedding, and denotes Pl\"ucker embedding. The two branches are encoded separately and fed to the diffusion transformer. In addition, to avoid requiring human-annotated trajectory preferences, we introduce trajectory preference optimization (TPO), a preference-learning scheme that keeps the video fixed and constructs negative examples by changing the conditioning trajectory, thereby isolating trajectory fidelity from video quality. After supervised fine-tuning (SFT), our 2B and 5B models reduce translation error by 49.4% and 36.8%, rotation error by 38.2% and 38.9%, FVD by 40.7% and 43.3%, and FID by 15.9% and 12.0%, respectively, relative to size-matched AC3D baselines. At the same total optimization-step count, SFT followed by TPO reduces the 2B model's translation and rotation errors by 4.4% and 4.0%, respectively, relative to SFT alone, while keeping FVD, FID, and CLIP score broadly stable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.