SAuBER: Neural Audio Coding via a Learnable MDCT Tokenizer and Energy-Based Refinement
Abstract
Generative neural audio codecs (NACs) achieve high perceptual quality at low bitrates, but typically rely on Generative Adversarial Networks (GANs) and spatial downsampling. These choices introduce training instability and limit autoencoder reconstruction fidelity. While hybrid approaches avoid spatial downsampling by pairing a fixed signal-processing transform with a learned compressor, designing them independently creates a domain mismatch. To resolve this, we propose SAuBER, a unified NAC framework that co-designs a learnable transform with a neural compressor while replacing adversarial generation with energy-based refinement. Specifically, our analysis/synthesis module operates as a learnable tokenizer built directly over patchified Modified Discrete Cosine Transform (MDCT) coefficients, producing near-transparent tokens without spatial downsampling. Within this token space, an information-bottleneck autoencoder extracts discrete latent representations, and an energy-based refinement head, trained via vector field matching, fits the gradient of an explicit energy function for stable, streamable generation. Autoencoder and refinement head are trained together to jointly optimize feature extraction, quantization, and refinement. At bitrates of 20–40 kbps, SAuBER achieves competitive performance against both classical signal-processing codecs and state-of-the-art GAN-based NACs in terms of perceptual quality and reconstruction fidelity.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.