acceptodds
Under review as a conference paper at ICLR 2027

Polarization-Aware Experts for Unified Video-Text-to-Audio Generation

Abstract

Unified video-text-to-audio generation supports Text-to-Audio (T2A), Video-to-Audio (V2A), and joint Video-Text-to-Audio (VT2A) within a single model. However, joint optimization introduces cross-task competition due to their heterogeneous conditioning requirements, as well as intra-task competition in VT2A, where excessive reliance on one modality can weaken alignment with the other. Existing systems primarily rely on shared dense backbones and progressive training, without explicitly preserving task-associated computation or directly regulating condition dominance in VT2A samples. We propose a two-stage Mixture-of-Experts framework that independently specializes T2A- and V2A-associated expert pools before composing them for unified multi-task training. Task-conditioned candidate restriction separates the computation paths of the single-condition tasks while allowing VT2A to access both pools. We further characterize excessive routing preference for one pool as expert-pool polarization and introduce a training-only intervention that reserves one routing slot for the under-selected pool while leaving the other unrestricted. Experiments demonstrate competitive T2A performance and strong V2A and VT2A results. Ablations validate the proposed specialization and candidate restriction, while the intervention improves targeted VT2A cases and mitigates inference-time expert-pool polarization. Our code will be publicly released.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.