Latent Spectral Equivariance
Abstract
Scalable audio generation systems typically cascade a pre-trained variational autoencoder (VAE) for audio compression and a latent diffusion model for distribution modeling. However, existing VAEs prioritize reconstruction fidelity without explicitly considering how readily the resulting latent distribution can be modeled by diffusion, potentially limiting generation quality. Recent work highlights the importance of aligning latent representations with the structure of the underlying data for diffusion modeling. Building on this direction, we introduce Latent Spectral Equivariance (LatentEQ), the first framework that systematically introduce equivariance into audio VAE optimization by pairing latent transformations with corresponding spectrogram transformations. Specifically, temporal scale equivariance applies matched low-pass filtering along time, aligning slowly varying latent components with slowly varying acoustic structure. Spectral granularity equivariance pairs latent channel prefixes with frequency-smoothed spectrograms, organizing channels to progressively capture spectral envelopes and fine harmonic detail. We instantiate LatentEQ on the VAEs of two well-known audio generation systems, Stable Audio Open and MMAudio, through a fine-tuning process without changing their architecture or introducing additional objectives. Comprehensive results demonstrate improved generation quality across waveform- and spectrogram-based VAEs and both text-to-audio and video-to-audio tasks, highlighting LatentEQ as an effective principle to learn structured audio latents.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.