FLORE-AE: Rethinking Audio Autoencoding with Waveform Diffusion Decoders
Abstract
Audio autoencoders for latent generation aim to combine compact representations with faithful and perceptually realistic waveform reconstruction. Existing spectral–adversarial decoders use learned discriminators, introducing minimax optimization. In this work, we instead formulate high-compression audio autoencoding as latent-conditioned waveform diffusion, where compact latents encode source structure and the diffusion decoder models remaining acoustic detail. FLORE-AE combines carrier–amplitude factorization for locally normalized synthesis, FLORE-DiT for time-aligned latent guidance and multiresolution context, and Endpoint-Aware MeanFlow for two-evaluation decoding of encoded and generated latents. Matched controls show that local factorization improves all reconstruction criteria, alignment drives paired fidelity, multiresolution computation matches full-resolution quality at lower inner-DiT cost, and endpoint-aware adaptation improves generated-latent decoding. Across speech, music, and sound, FLORE-AE improves reconstruction over the evaluated spectral–adversarial and generative autoencoders while supporting generation and editing at high compression. These results demonstrate a practical non-adversarial waveform decoder for compact audio representations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.