Spend-Where-It-Matters (SWIM): Content-Aware Efficient Diffusion Models for Image and Video Generation
Abstract
Diffusion models achieve state-of-the-art performance in image synthesis and editing, but remain costly due to long denoising trajectories and large backbone models. Existing accelerations reduce the number of sampling steps or reuse cached features on a static schedule, so every input receives the same computation, even though which parts of the backbone matter varies across inputs and timesteps. We introduce SWIM, a model-agnostic, content-aware diffusion framework that decides, at each denoising step, which backbone blocks to execute under a user-set compute budget, while leaving the backbone weights frozen. A lightweight recurrent gating module reads the current sample's intermediate features, the change in the predicted score across diffusion steps, and bypasses the blocks that contribute least to the current update, so compute is spent where the sample needs it rather than on a fixed schedule. Because per-step compute is no longer uniform, we report cumulative multiply–accumulate operations (CMACs) as the measure of compute cost, over the full trajectory. For SD3.5 T2I generation, SWIM uses under 50% of DeepCache's CMACs yet outperforms it across all T2I-CompBench++ benchmarks with a 38% relative mean score improvement. On FLUX 2, SWIM cuts CMACs by 23% while staying within 0.01 of the T2I-CompBench++ mean. On ImageNet-256, SWIM + PixelDiT uses 18% fewer CMACs and 19% fewer parameters than PixelDiT at competitive FID, and improves FID over DiT-XL with 79% fewer CMACs. SWIM also extends to video generation, reducing CogVideoX-5B inference time by 13%.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.