Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction
Abstract
Continuous-latent audio autoencoders form the backbone of latent music generators, and their reconstruction fidelity limits the acoustic detail available to downstream generators. At the compression rates these pipelines require, decoders commonly exhibit three recurring failure modes: high-frequency loss, phase incoherence, and stereo-image collapse. These shortcomings in high-fidelity generation share a structural root: waveform autoencoders, previously regarded as the frontier design, lack an explicit frequency axis, leaving no handle for targeted per-band correction. Among five matched-budget representations in our controlled study, the complex STFT achieves the lowest full-band and high-frequency spectral distances, providing direct access to magnitude and phase at every frequency bin. Building on this finding, we present ear-VAE2 , a complex-spectral autoencoder that exploits the explicit frequency structure and introduces cross-channel interaction between the left and right channels. First, Spec-SnakeBeta exploits the additional capacity of the frequency-domain representation by learning a periodic activation for each frequency bin, with frequency-dependent initialization in log-parameter space. This design overall outperforms the other activation variants in our controlled ablation, while using fewer parameters than the fully independent variant. Second, Duplex-Aware Refiner applies band-specific corrections to magnitude and phase. Its band allocation follows duplex theory of sound localization and corrects the decoder's reconstruction error. On the 546-track Song Describer Dataset, ear-VAE2 achieves the best point estimates on five of seven reconstruction metrics and matches the best stereo-coherence score. Adding the Duplex-Aware Refiner module further reduces Mel Distance by 19.4%, and its duplex band allocation uses ≈45% fewer residual-output dimensions than the Unconstrained Refiner version. Beyond this efficiency, the banded allocation also lowers spectral distances and targeted spatial-cue error metrics, and it receives higher mean paired ratings from professional mixing and mastering engineers. The downstream generator using ear-VAE2 latents achieves better point estimates on all 12 automatic downstream metrics. Demo page is available at: https://earvae2-2026.github.io/.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.