Follow the Ball: Do Video Diffusion Models Control Where An Object Moves?
Abstract
Video diffusion models can generate realistic futures from short input videos, yet they remain unreliable at following instructions about how objects should move. Despite rapid progress in controllable video generation, it remains unclear whether these instructions actually guide the generated motion or whether models mainly extrapolate from the input video. We begin by asking how much a generated object’s motion changes when its instructed path changes. To separate instruction following from general future prediction, we consider two complementary tests: (1) response to intervention, where two generations share the same input video and randomness but receive different object paths, and (2) future similarity, where each generation is compared with the recorded continuation. Broadcast sports provide a natural testbed for this distinction because the movement of a ball or shuttlecock is central to what happens next. Using this setting, we introduce SportWorldBench, a controlled framework for evaluating object-motion control in video generation. Across changes to path labels, how paths are provided, model training, video resolution, and random samples, we find no reliable response to changed paths. Moreover, at higher resolution the models resemble the recorded future more closely than simply repeating the last input frame, yet this resemblance is nearly identical whether the instructed path matches the recording or is deliberately altered. These findings show that generating a plausible future does not imply following object-motion instructions and motivate methods that make such controls reliably influence generation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.