Temporal-Attention Head Specialization During Video Diffusion Training
Abstract
Video diffusion transformers depend on temporal attention to coordinate information across frames, yet existing analyses of this mechanism examine only trained models. Such final-checkpoint analyses show what temporal structure exists but not when or where it forms during training, which is the information needed to monitor training and to target head-level interventions such as pruning. Population averages can also hide this structure, because the increase in a few heads and the decrease in most other heads offset each other in the mean. We therefore conduct a checkpoint-resolved census of every temporal-attention head across nine Open-Sora STDiT training runs spanning three model scales (306M to 1.03B parameters). Each head is scored by its cross-frame attention concentration (CFAC), the entropy-normalized concentration of its attention over the 16 frames. CFAC is 0 when a head spreads attention evenly across all frames and approaches 1 when it routes each frame's attention to a few specific frames. Heads are then selected under a preregistered change-point and effect-size rule. Per-head analysis shows that a small minority of heads, roughly 4-13% in full-grid runs, becomes strongly concentrated during training, while aggregate CFAC stays flat or decreases in every run. Across seeds, selected heads repeatedly appear in the first temporal block, but the specific heads selected within that block differ across seeds. In the 760M runs, selected heads show two main recurring frame-to-frame attention patterns. In the first, a self-frame diagonal, each frame attends mostly to itself. In the second, an adjacent-frame band, each frame attends mostly to its neighbors. These patterns recur across runs even when they emerge in different heads (i.e., at different layer and head indices). Our correlation and ablation experiments do not establish that these heads affect generated video quality. The census nevertheless identifies when and where temporal specialization forms during training. The same checkpoint-resolved, per-head analysis under fixed selection rules can provide this information for other factorized video diffusion transformers and, with adapted routing metrics, for joint spatio-temporal architectures.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.