EcoVideo++: Characterizing and Exploiting Tri-dimensional Capability Boundaries for Collaborative Video Diffusion
Abstract
Diffusion Transformers (DiTs) are widely used for video generation, but their high computational cost limits deployment. Existing collaborative methods switch between large and small models at different denoising timesteps. However, this timestep-level strategy overlooks the multi-dimensional nature of the capability gap. We show that the gap varies across timesteps, concentrates on difficult frames, and is further localized to a small number of tokens. Thus, applying the large model to an entire timestep is unnecessarily coarse-grained. Based on these observations, we propose EcoVideo-β, a collaborative framework that coordinates large and small models at the timestep, frame, and token levels. The small model generates the full video to maintain global consistency, while the large model selectively refines challenging spatiotemporal regions. A pipelined edge–cloud execution strategy further reduces collaboration overhead. EcoVideo-β enables fine-grained capability allocation for efficient, high-quality video generation.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.