acceptodds
Under review as a conference paper at ICLR 2027

EcoVideo++: Characterizing and Exploiting Tri-dimensional Capability Boundaries for Collaborative Video Diffusion

Abstract

Diffusion Transformers (DiTs) are widely used for video generation, but their high computational cost limits deployment. Existing collaborative methods switch between large and small models at different denoising timesteps. However, this timestep-level strategy overlooks the multi-dimensional nature of the capability gap. We show that the gap varies across timesteps, concentrates on difficult frames, and is further localized to a small number of tokens. Thus, applying the large model to an entire timestep is unnecessarily coarse-grained. Based on these observations, we propose EcoVideo-β, a collaborative framework that coordinates large and small models at the timestep, frame, and token levels. The small model generates the full video to maintain global consistency, while the large model selectively refines challenging spatiotemporal regions. A pipelined edge–cloud execution strategy further reduces collaboration overhead. EcoVideo-β enables fine-grained capability allocation for efficient, high-quality video generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.