acceptodds
Under review as a conference paper at ICLR 2027

Heavy-First, Light-Late: Reversing Capacity Scheduling for Efficient Video Diffusion

Abstract

Diffusion transformers achieve state-of-the-art video generation quality but incur prohibitive inference latency due to repeated full-capacity denoiser evaluations across sampling steps. Many multi-capacity methods follow a coarse-to-fine paradigm, allocating higher computational budgets to late-stage refinement. In contrast, we show that capacity demands peak during early denoising: We formalize this property as , showing that executing a brief prefix of early steps with a large model locks the generative path, after which a compact model can complete the remaining steps with minimal fidelity loss. Motivated by this observation, we propose (daptive etwork apacity andoff for ptimal outing), a training-free cross-scale inference framework that dynamically routes computation between heterogeneous diffusion transformers. ANCHOR introduces a coordinate alignment mechanism to reconcile latent representations across model scales without fine-tuning, alongside an adaptive handoff predictor that determines sample-specific transition points from early trajectory dynamics. Comprehensive evaluations on Wan2.1, CogVideoX, and SkyReels-V2 demonstrate up to speedup while preserving high frame-aligned similarity to the full-Large reference. Furthermore, our adaptive routing exposes superior fidelity-speed Pareto operating points by dynamically assigning longer prefixes to complex video samples.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.