acceptodds
Under review as a conference paper at ICLR 2027

UniCaMo: Unified Camera and Motion Control for Video Diffusion Models

Abstract

We present UniCaMo, an unified framework for joint control of camera and object motion in image-and-text-to-video diffusion models. Given a reference image, text prompt, sparse 3D object trajectories, and a 3D camera trajectory, UniCaMo constructs motion-consistent input noise that directly encodes the desired motion. Our 3D-grounded noise construction combines reference-derived noise propagation with a spherical Gaussian noise representation for newly revealed regions, providing a simple and effective motion prior without requiring reconstruction of unseen scene geometry. UniCaMo requires no architectural modifications or dedicated control modules and can be adapted using lightweight LoRA fine-tuning with low training data volume. Experiments on standard controllable video generation benchmarks demonstrate state-of-the-art motion controllability and high visual quality with minimal inference overhead. Qualitative videos are presented on the HTML webpage provided in the submitted supplementary zip files.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.