Efficient Conditional Modeling for Controllable Video Generation
Abstract
Despite substantial recent progress in controllable video generation with various conditioning signals, current frameworks suffer from severe computational inefficiency when applied to video diffusion transformers. The issue becomes particularly acute when the conditioning signal itself is a video, resulting in excessively long contextual sequences. To address this limitation, we propose EasyVideoControl, an efficient framework for conditional video generation with DiTs (diffusion transformers). EasyVideoControl introduces two core components. First, we present the Asymmetric Shared Causal Mixture-of-Transformers structure, which incorporates a frozen copy of the base diffusion transformer that shares self-attention layers with the backbone under a causal attention scheme. By training only low-rank adaptation (LoRA) parameters in this unidirectional conditioning branch, the model maximizes reuse of pretrained knowledge while minimizing training computational overhead. This design enables the backbone to be modulated by the conditional branch without reciprocal dependency, supporting KV Cache for efficient inference. Second, we introduce Reconstructive Conditioning Compression and Recovery, a lightweight encoder–decoder mechanism that compresses and reconstructs conditional tokens. This approach greatly reduces context length while preserving fine-grained conditional details through a reconstruction-based objective. Extensive experiments across multiple conditional video generation tasks show that EasyVideoControl achieves comparable or superior performance while being significantly more efficient than existing video control methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.