Building Audio Scenes with Voice Design and Cloning
Abstract
Building audio scenes requires jointly generating multi-speaker speech, ambience, music, and sound effects. This task requires consistent and expressive speech coordinated with the surrounding audio. Speech can be generated through voice design from a scene description or through voice cloning from reference audio. However, the scarcity of high-quality annotated data and the difficulty of supporting both tasks while jointly generating multiple audio modalities further complicate audio scene generation. Therefore, we first propose VoxCaps, which cleans media data, expands it with targeted synthetic data, and uses a persona library to annotate diverse multi-level scene descriptions. Then, we introduce VoxVAE, which pairs a local anti-aliased encoder with a high-capacity decoder to provide a shared continuous representation for multiple audio modalities. Next, we propose VoxDirector, an expressive multi-speaker speech and surrounding audio generation model that supports both voice design and cloning. VoxDirector adopts reward-conditioned quality control to improve generation quality, and a Dual-Router MoE tailored to the two tasks and multiple audio modalities. Experimental results show that VoxDirector outperforms the baselines on key metrics for both voice design and cloning, while supporting the generation of complex audio scenes.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.