Spatial Audio Generation and Editing with Executable Plans and Self-Distilling Transfusion
Abstract
Generating and editing multi-source spatial audio requires fine-grained control over source semantics, temporal activity, and 3D trajectories, while preserving unmodified acoustics during edits. Unifying these tasks requires explicit source-level instruction grounding and preservation of multichannel spatial relationships. We present AMBIT, an end-to-end framework for native First-Order Ambisonics (FOA) generation and editing using an executable ScenePlan that binds source descriptions to activity intervals and 3D trajectories. Built on a unified Transfusion backbone, our framework introduces three contributions: (1) autoregressive planning coupled with rectified-flow rendering, using contrastive reference features for instruction grounding and full reference latents for acoustic preservation; (2) an FOA-aware continuous VAE with shared-decoder omnidirectional recovery, asymmetric grouped KL, and spatial covariance matching; and (3) on-policy self-distillation (OPSD) tailored to Transfusion, jointly refining planning and rendering with target-conditioned teacher predictions on student-generated prefixes and sampling states. Target audio is used only for teacher conditioning during refinement, not inference. Evaluations across speech, music, and sound scenes show improved generation content and spatial quality over the compared baselines, alongside better acoustic fidelity and unedited-source preservation during targeted edits.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.