Unified Spatial Audio Generation via Iterative Expert Routing
Abstract
Spatial audio generation aims to synthesize spatial sound fields that faithfully reflect user intent under diverse conditioning signals. However, existing methods remain limited in their ability to generate accurate, high-fidelity spatial audio from heterogeneous inputs. To address this limitation, we propose UniAmbi, a unified framework for jointly synthesizing environmental audio and speech in first-order Ambisonics (FOA) format from multi-conditions. Our framework integrates three complementary designs to achieve efficient generation with semantic consistency, spatial accuracy, and perceptual fidelity. First, we introduce a Looped DiT backbone that addresses the complexity of multi-conditional FOA generation through iterative computation with shared parameters, increasing effective depth without a proportional increase in parameter count. Second, we introduce a loop aware sliding window Mixture-of-Experts (MoE) module to mitigate the representational limitations of strict parameter sharing. By incorporating expert routing on the current loop iteration and local temporal context, the module enables round-dependent transformations while encouraging consistent expert selection among neighboring tokens. % Finally, we introduce per-loop on-policy self-distillation (OPSD) to supervise predictions at every loop depth along the model’s own generation trajectories, using feedback on spatial accuracy, semantic consistency, and perceptual quality. Finally, we introduce per-loop on-policy self-distillation (OPSD), which converts generation-level feedback into direct supervision for predictions at every loop depth along the model's own generation trajectories. Experimental results show that our model outperforms existing methods on multiple spatial audio generation tasks. Demos can be found at https://uniambi.github.io.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.