acceptodds
Under review as a conference paper at ICLR 2027

Learning Acoustic Scene Decomposition

Abstract

Acoustic experiences are formed not only by sound content but also by propagation effects characterized by room impulse responses (RIRs). We study the problem of decomposing a single audio recording into the latent factors that produced it: the dry audio of each source and the room impulse response (RIR) through which it reaches the listener. Existing methods tackle pieces of this problem in isolation, focusing on RIR estimation, audio source separation, or dereverberation, typically under single-source assumptions. We instead pose it as a single inverse problem and present ASCEND, a diffusion-based framework that recovers all components jointly. Our key idea is that a correct decomposition, when re-rendered through the known acoustic forward model, must reproduce the input mixture. We turn this into an inference-time signal via consistency-guided sampling: at every denoising step, we re-render the current estimates, measure the mismatch with the input, and use its gradient to guide sampling. The denoiser is a transformer model with attention that operates both within and across audio source streams. On large-scale simulated benchmarks, ASCEND outperforms task-specific baselines on RIR estimation and dry-source extraction, with gains that transfer zero-shot to real-world recordings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.