SoundscapeGen: Smart Spatial Soundscape Construction from High-level Descriptions
Abstract
Generating spatial soundscapes from high-level natural-language descriptions requires not only realistic audio synthesis but also reliable control over what sound events occur, when they occur, and how they move in three-dimensional space. Existing end-to-end and compositional approaches often rely on the implicit knowledge of large language models (LLMs) for scene planning, and provide limited guarantees that the generated audio follows the intended semantic, temporal, and spatial structure. We introduce SoundscapeGen, a framework that represents a soundscape as a set of associated controllable sound objects and constructs it through a Plan → Generate → Verify → Compose pipeline. An acoustic scene knowledge graph (ASKG) provides explicit priors for event selection, temporal relations, source properties, and motion, while a fine-tuned physics-guided First-Order Ambisonics generator synthesizes each sound object with explicit trajectory control. Object-level semantic, temporal, and spatial verification further filters inconsistent generations before composition. Experiments show that SoundscapeGen improves event coverage and temporal consistency over both end-to-end text-to-audio and compositional baselines, while providing explicit control over individual source trajectories. Its object-based representation further supports localized interactive editing without regenerating the entire soundscape. Demos are available at https://soundscapegeneration.github.io/SoundscapeGen/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.