Reading Out Camera-Decoupled Motion from Frozen Video Diffusion Features
Abstract
Understanding scene motion independently of camera movement is fundamental to robotics, autonomous navigation, and video understanding. Apparent motion in monocular video, however, entangles scene dynamics with camera movement. Video diffusion models generate coherent object motion under diverse camera movement, raising the question of whether their frozen features support camera-decoupled motion readout. We probe this question by keeping the video diffusion backbone frozen and training only a lightweight decoder for moving object segmentation (MOS) and motion magnitude estimation (MME). MOS captures the clip-level extent of independently moving objects, while MME estimates the per-pixel, scene-scale-normalized magnitude of world-space 3D displacement between consecutive frames. These readouts are strong enough to outperform dedicated moving object segmentation methods, while comparing favorably with leading feed-forward 4D reconstruction models in 3D motion magnitude estimation. Through controlled comparisons with geometric features, including synthetic-to-real transfer, we provide evidence that this capability is rooted in pretrained video diffusion features rather than decoder training alone. Together, these results establish frozen video diffusion features as effective representations for reading out camera-decoupled scene motion at both object and point levels.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.