The Seriality Gap in Video Diffusion Models
Abstract
When one ball strikes another, then another, video models should predict the consequences of each bounce. In controlled experiments on multi-ball hard-sphere dynamics, we find that the performance of standard bidirectional video diffusion degrades as the causal chain lengthens, even when provided more denoising steps. In a length-matched control, where ball–ball interactions are absent, the degradation largely disappears, isolating dependent-event structure rather than video length as the cause. Across intervention studies, **methods that increase effective serial computation improve performance disproportionately**, including autoregressive/blockwise generation and architectural depth. This finding extends to both video prediction and simulator-state prediction, and also applies to both physical and algorithmic tasks. We identify this pattern as the **seriality gap**: a mismatch between tasks requiring growing serial computation and video diffusion models whose denoising loop does not provide scalable serial compute. We then prove that, for fully observed deterministic video prediction with an exact score, denoising steps do not add serial computation beyond the backbone, indicating a structural obstacle for video diffusion on serial reasoning and simulation tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.