EVA: Efficient Training for Joint Video-Audio Generation via Selective Cross-Modal Interaction
Abstract
Joint video–audio generation has advanced rapidly, with most recent approaches building upon pretrained text-to-video and text-to-audio models to construct joint generators. However, dense cross-modal interaction throughout training can introduce redundant or noisy signals in randomly initialized interaction modules, impairing both training efficiency and generation quality. Thus, we present an Efficient training framework for joint Video-Audio generation, namely EVA, for fast adaptation from pretrained uni-modal generators through multi-granularity selective cross-modal interaction: (1) at the task level, EVA selectively enables audio-to-video, video-to-audio, or bidirectional interaction according to the dominant cross-modal dependency of each sample; (2) at the model level, it restricts cross-modal interaction to deeper layers to reduce unnecessary computation and avoid premature interference with uni-modal representations; (3) at the token level, adaptive gating modulates cross-modal attention to emphasize informative visual–audio correspondences and ignore irrelative interactions, improving synchronization while preserving uni-modal generation quality. On Verse-Bench, our 7B model achieves competitive visual, audio, and speech quality with substantially lower training cost. It also obtains the lowest DeSync of 0.1902 and the highest LSE-C of 7.8018, demonstrating strong audio–visual synchronization. By reducing the cost of cross-modal adaptation, EVA can free training budget for scaling model capacity or further optimization. Code and weights will be released soon.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.