acceptodds
Under review as a conference paper at ICLR 2027

Latent Spectral Equivariance

Abstract

Scalable audio generation systems typically cascade a pre-trained variational autoencoder (VAE) for audio compression and a latent diffusion model for distribution modeling. However, existing VAEs prioritize reconstruction fidelity without explicitly considering how readily the resulting latent distribution can be modeled by diffusion, potentially limiting generation quality. Recent work highlights the importance of aligning latent representations with the structure of the underlying data for diffusion modeling. Building on this direction, we introduce Latent Spectral Equivariance (LatentEQ), the first framework that systematically introduce equivariance into audio VAE optimization by pairing latent transformations with corresponding spectrogram transformations. Specifically, temporal scale equivariance applies matched low-pass filtering along time, aligning slowly varying latent components with slowly varying acoustic structure. Spectral granularity equivariance pairs latent channel prefixes with frequency-smoothed spectrograms, organizing channels to progressively capture spectral envelopes and fine harmonic detail. We instantiate LatentEQ on the VAEs of two well-known audio generation systems, Stable Audio Open and MMAudio, through a fine-tuning process without changing their architecture or introducing additional objectives. Comprehensive results demonstrate improved generation quality across waveform- and spectrogram-based VAEs and both text-to-audio and video-to-audio tasks, highlighting LatentEQ as an effective principle to learn structured audio latents.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.