acceptodds
Under review as a conference paper at ICLR 2027

SEAE: Spectral Energy and Ablation Error for Structural Pruning of Few-step Video Diffusion Transformers

Abstract

The increasing size of Video Diffusion Models (VDMs) makes inference computationally expensive even after step distillation. Few-step distillation accelerates VDMs by reducing the denoising steps, but each step still executes the full width attention modules and the full depth Diffusion Transformer (DiT), leaving substantial memory usage and computation costs. Few-step VDMs Structural pruning is challenging because each solver step performs a larger portion of the noise-to-video transformation, giving pruning errors fewer subsequent steps in which to be corrected. We introduce SEAE, a structural pruning framework that reduces both attention width and DiT depth in few-step VDMs. Specifically, (i) our MHA-CSS method constructs Gram matrices from projected attention-head activations at all solver timesteps, uses their cumulative spectral energy to determine how many heads to retain in each self- and cross-attention module, and then selects the retained heads through CSS-inspired backward elimination with scalar refitting; (ii) we score each DiT blocks using Ablation Error, which temporarily removes each DiT block at a given timestep, completes the remaining sampling process, and measures its influence by comparing the final generated latent with the unablated reference. We form a low error candidate set at each solver timestep and retain their intersection as one static block pruning mask. Both masks are derived solely from forward passes over calibration set and define a fixed compact architecture shared across all timesteps, without runtime routing or cross-timestep feature caching. Recovery distillation finally trains the compact student using a weighted combination of teacher-student distillation and Flow Matching losses. Our model reduces the parameter count by 25.4% relative to the four-step Wan2.2-TI2V-5B teacher model, while nearly preserving its VBench-I2V total score for video generation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.