What Controls Motion in Video Diffusion Models? A Mechanistic Study
Abstract
Motion is central to video generation models, yet what inside the model controls it remains largely unexplored. Some methods control object motion by training an added module on video data. Others require no training and edit the latent or the attention at inference time. These methods change the generated motion, but they do not show what inside the model produces the change. We present a mechanistic study of that question. We train a video diffusion transformer on synthetic videos, where the launch angle of a ball is controllable, and we read that angle in degrees from generated videos. This model allows us to control object motion more precisely than a pretrained model does. We present three findings. First, adding one shared vector at every token position, the standard form of activation steering, does not rotate the trajectory. Second, letting the added vector vary across positions does reach the request. Third, the control depends on a small group of conditioning-related attention heads, and holding them at their un-intervened outputs removes most of the rotation. The same pattern holds on the pretrained Wan2.1-1.3B. Together these experiments identify what controls motion inside video diffusion models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.