Universal Mixup Loss: Training Audio Autoencoders to Decode the Positive-Weight Mixing Cone
Abstract
Waveform mixing is linear, but mixing the latent codes of an audio autoencoder rarely produces a representation that decodes to the corresponding audio mixture. This limits the use of latent arithmetic for audio editing, source separation, and generative mixing. We show that supervising a few gain and addition identities leaves large regions of the mixing space unconstrained. We introduce the Universal Mixup Loss, a decoder-side objective that instead supervises positive-weight combinations of two latent codes throughout their mixing cone. Training points are sampled in polar coordinates, decoupling an annealed radial scale from a uniformly distributed mixing angle. Paired with a scale-sensitive soft-norm bottleneck, this supervision substantially reduces mixing commutation errors while preserving reconstruction quality. The additive noise injected at the bottleneck improves tolerance to the residual codes produced by subtraction and supports decoding at extreme gains. As a result, latent subtraction can recover a source from a mixture when the codes of the interfering sources are available, and decoding remains consistent across interpolation laws absent from training. The resulting latent structure also benefits conditional rectified flow: its velocity field is easier to regress, arithmetic remains effective on generated codes, and high-quality outputs are reached with fewer sampling steps.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.