Rethinking Mixture-of-Experts for Audio-Visual Diffusion Models
Abstract
This work systematically investigates how to design and train Mixture-of-Experts (MoE) architectures for large-scale audio-visual diffusion models. Audio-visual latent tokens form continuous representations with strong spatial, temporal, and cross-modal correlations, motivating a re-evaluation of existing MoE design choices. Through systematic study, we summarize sparse capacity design principles tailored to audio-visual generation. We strengthen text conditioning with richer semantic representations while keeping the text conditioning branch dense and allocating sparse capacity to the visual backbone. Meanwhile, we find that load variation persists under continuous balancing and can arise from either semantic specialization or optimization instability. We therefore combine routing regulation with targeted stabilization to preserve normal expert differentiation while suppressing instability-driven collapse. Finally, we validate these designs using millions of high-quality audio-visual samples in a large-scale distributed training environment, establishing architectural and training principles for large-scale audio-visual diffusion MoE models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.