TCMU: Temporal Camera Motion Understanding
Abstract
Camera motion is temporal and compositional: multiple operations can overlap, change speed, and transition within a single shot, whereas shot-level labels leave this structure unspecified. We introduce Temporal Camera Motion Understanding (TCMU), a task that predicts a sequence of temporal segments specifying the types, directions, and speeds of all active camera motions. To support training and evaluation, we construct approximately 292K human-annotated video clips and TCMU-Bench, an independent benchmark of 1,000 shots with 2,590 manually annotated segments. We develop two approaches. TCMU-Geo derives motion timelines from reconstructed geometry by separating depth-independent and depth-dependent components of camera-induced image motion. TCMU-VLM learns to generate these timelines directly from video using the annotated training dataset. On TCMU-Bench, TCMU-Geo achieves 60.7 frame F1, compared with 32.8 for the pose-based SpatialVID baseline. TCMU-VLM achieves 71.5 frame F1, compared with 25.4 for the strongest evaluated zero-shot VLM, as well as 70.0 Event-F1 at temporal IoU 0.5 and 84.9 Boundary-F1 at a 0.5-second tolerance. TCMU-VLM is stronger overall in recognition and temporal localization, while TCMU-Geo performs better on parallax-based motion discrimination and directed-motion recall under dense concurrency. Paired event analysis reveals complementary motion intervals recovered by only one method. As an additional application, geometry-assisted human review of training annotations improves VLM recognition and event localization, although boundary-localization accuracy decreases. We will release TCMU-Bench, the evaluation code, the distilled TCMU-VLM-2B checkpoint, and the complete TCMU-Geo code.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.