AuDiAE: Learning Audio Latents by Coupling Diffusion Priors and Decoders
Abstract
Audio autoencoders determine both the information available for waveform reconstruction and the latent target learned by a downstream generator. Standard Gaussian posterior-KL regularization limits latent information but also penalizes mismatch between the aggregate posterior and a fixed Gaussian. Under a tight KL budget, this additional distribution-matching cost leaves less capacity for reconstruction-relevant detail. Downstream conditional generation additionally depends on how much latent variation is predictable from a condition such as text, while high compression leaves recording-specific acoustics underdetermined for deterministic waveform decoders. We introduce AuDiAE, an Audio Diffusion AutoEncoder that couples a learned diffusion rate model with a generative waveform decoder through a common encoder representation. The marginal diffusion prior replaces the fixed reference with a learned density while maintaining a variational information bound. Its conditional counterpart makes condition-predictable information less costly in the training objective while the marginal rate continues to control total latent information. A flow-matching decoder trained to predict the clean waveform models residual variation and values encoded source detail through its reduction in optimal denoising risk. We evaluate AuDiAE at 2048× and its high-compression variant, AuDiAE-H, at 4096× on reconstruction and matched AudioCaps generation. AuDiAE leads every reported reconstruction metric and improves the strongest external AudioCaps FAD by 35% while attaining the highest CLAP. Matched controls support the learned rate reference and correct text–audio pairing; frozen diagnostics isolate tighter prior fit and text use without supporting a causal decoder advantage.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.