Predictability in Video Diffusion: Early Attention Predicts Visual Failures
Abstract
Text-to-video diffusion models require costly multi-step denoising, yet generation failures are usually identified only after the final video is complete. We find that intermediate video-to-text attention contains early information about whether individual visual constraints will be satisfied. Motivated by this insight, we introduce VideoProbe, a lightweight spatiotemporal probe that predicts token-level failures using only intermediate cross-attention and the denoising timestep. We also construct AttnPulse-10K, a benchmark that aligns multi-step attention from 10,000 Wan2.2 trajectories with uncertainty-masked token supervision derived from the completed videos. On Wan2.2, a single probe trained across nine denoising steps achieves mean AUPRC and generalizes without retraining to unseen denoising steps, including the first denoiser forward ( AUPRC). Shuffling attention across samples reduces AUPRC to , supporting the conclusion that matched attention carries predictive information. Independent experiments on Wan2.1 and HunyuanVideo-1.5 further show that early attention predicts final failures across different video generators and attention architectures. These early token-level predictions could support seed routing, prompt refinement alerts, and RL rollout mining.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.