Polarization-Aware Experts for Unified Video-Text-to-Audio Generation
Abstract
Unified video-text-to-audio generation supports Text-to-Audio (T2A), Video-to-Audio (V2A), and joint Video-Text-to-Audio (VT2A) within a single model. However, joint optimization introduces cross-task competition due to their heterogeneous conditioning requirements, as well as intra-task competition in VT2A, where excessive reliance on one modality can weaken alignment with the other. Existing systems primarily rely on shared dense backbones and progressive training, without explicitly preserving task-associated computation or directly regulating condition dominance in VT2A samples. We propose a two-stage Mixture-of-Experts framework that independently specializes T2A- and V2A-associated expert pools before composing them for unified multi-task training. Task-conditioned candidate restriction separates the computation paths of the single-condition tasks while allowing VT2A to access both pools. We further characterize excessive routing preference for one pool as expert-pool polarization and introduce a training-only intervention that reserves one routing slot for the under-selected pool while leaving the other unrestricted. Experiments demonstrate competitive T2A performance and strong V2A and VT2A results. Ablations validate the proposed specialization and candidate restriction, while the intervention improves targeted VT2A cases and mitigates inference-time expert-pool polarization. Our code will be publicly released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.