acceptodds
Under review as a conference paper at ICLR 2027

Efficient Conditional Modeling for Controllable Video Generation

Abstract

Despite substantial recent progress in controllable video generation with various conditioning signals, current frameworks suffer from severe computational inefficiency when applied to video diffusion transformers. The issue becomes particularly acute when the conditioning signal itself is a video, resulting in excessively long contextual sequences. To address this limitation, we propose EasyVideoControl, an efficient framework for conditional video generation with DiTs (diffusion transformers). EasyVideoControl introduces two core components. First, we present the Asymmetric Shared Causal Mixture-of-Transformers structure, which incorporates a frozen copy of the base diffusion transformer that shares self-attention layers with the backbone under a causal attention scheme. By training only low-rank adaptation (LoRA) parameters in this unidirectional conditioning branch, the model maximizes reuse of pretrained knowledge while minimizing training computational overhead. This design enables the backbone to be modulated by the conditional branch without reciprocal dependency, supporting KV Cache for efficient inference. Second, we introduce Reconstructive Conditioning Compression and Recovery, a lightweight encoder–decoder mechanism that compresses and reconstructs conditional tokens. This approach greatly reduces context length while preserving fine-grained conditional details through a reconstruction-based objective. Extensive experiments across multiple conditional video generation tasks show that EasyVideoControl achieves comparable or superior performance while being significantly more efficient than existing video control methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.