acceptodds
Under review as a conference paper at ICLR 2027

SAuBER: Neural Audio Coding via a Learnable MDCT Tokenizer and Energy-Based Refinement

Abstract

Generative neural audio codecs (NACs) achieve high perceptual quality at low bitrates, but typically rely on Generative Adversarial Networks (GANs) and spatial downsampling. These choices introduce training instability and limit autoencoder reconstruction fidelity. While hybrid approaches avoid spatial downsampling by pairing a fixed signal-processing transform with a learned compressor, designing them independently creates a domain mismatch. To resolve this, we propose SAuBER, a unified NAC framework that co-designs a learnable transform with a neural compressor while replacing adversarial generation with energy-based refinement. Specifically, our analysis/synthesis module operates as a learnable tokenizer built directly over patchified Modified Discrete Cosine Transform (MDCT) coefficients, producing near-transparent tokens without spatial downsampling. Within this token space, an information-bottleneck autoencoder extracts discrete latent representations, and an energy-based refinement head, trained via vector field matching, fits the gradient of an explicit energy function for stable, streamable generation. Autoencoder and refinement head are trained together to jointly optimize feature extraction, quantization, and refinement. At bitrates of 20–40 kbps, SAuBER achieves competitive performance against both classical signal-processing codecs and state-of-the-art GAN-based NACs in terms of perceptual quality and reconstruction fidelity.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.