acceptodds
Under review as a conference paper at ICLR 2027

Self-Supervised Learning of Structured Motion from Videos

Abstract

Understanding motion in video is a fundamental challenge for visual learning, as frame-to-frame change entangles two sources of motion: camera and object. This decomposition has remained underexplored in representation learning, partly because these factors are tightly coupled in natural videos and difficult to supervise separately. Yet recovering it is important for learning robust motion representations that separate meaningful object movement from camera-induced variation. We study whether such structured motion representations can be recovered from frozen features of a pretrained image vision transformer. We propose the Structured Motion Model (SMM), which explicitly separates the dominant source of temporal change from residual motion through feature prediction, rather than representing video change with a single entangled latent or with unstructured, spatially dense transition tokens. Training combines self-supervised learning on real video with weak supervision of scene motion on synthetic Kubric data. We evaluate SMM on ProbeMotion, a new evaluation suite spanning synthetic and real videos with camera, object, and combined motion. SMM outperforms backbone baselines using global CLS or average-pooled features, and compares favorably to strongly supervised representations such as VGGT on several probes, despite using substantially weaker supervision. These results suggest that pretrained image models can be readily repurposed into structured motion representations, providing a useful inductive bias for learning and analyzing latent video motion.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.