Principal Component Alignment for Training a General Full-Band Stereo Audio VAE
Abstract
Latent alignment to a frozen foundation model is now standard practice for audio VAEs. The two choices it involves, which teacher and how its features are projected, turn out to govern diffusability: how well a diffusion model trained on the resulting latent space actually generates. We revisit both. Where prior work learns the projection relating teacher features to the latent, we align to a principal-component basis of those features. As the teacher we use the audio encoder of a perceptual quality predictor rather than a semantic encoder, and for general audio this proves the better choice. Side-by-side listening tests confirm improved prompt following and audio quality in the generations. Reconstruction quality, meanwhile, turns out to be an unreliable guide: the configuration that reconstructs better is not the one that generates better. The improved diffusability then buys room in the latent to carry a second channel at no cost. Our 128-dimensional stereo latent ties its otherwise identical 64-dimensional mono counterpart at 50/50 on prompt following and on downmixed audio quality. The result is a full-band 48 kHz VAE for general audio (speech, music, other sounds and mixtures of these) at a 50 Hz latent rate, which wins over 70% of audio-quality comparisons against the audio VAEs of LTX 2.5 and Stable Audio 3 under a fixed diffusion transformer and training budget.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.