Structuring Spatial Audio Representations via Invariance-Equivariance Learning
Abstract
A spatial audio representation must jointly capture scene semantics to recognize what is sounding and scene geometry to understand where sound events occur. Self-supervised objectives often favor semantic invariance, treating direction as a nuisance, making spatial geometry difficult to preserve. We propose a framework for structuring spatial audio representations through invariance–equivariance learning, using geometric transformations as a training signal. In first-order ambisonics (FOA), rotating a sound field is a matrix operation, providing rotated views and their relative rotations without labels. We train encoders to remain stable under spectral nuisances while predicting how their representations change under rotation. Within this framework, we compare representations explicitly partitioned into invariant and equivariant subspaces with an unsplit predictive latent. The latter, built on seq-JEPA, is to our knowledge the first joint-embedding predictive architecture (JEPA) for spatial audio that learns invariance and equivariance jointly. We show that this latent representation organizes direction as an ordered spherical map while preserving linearly separable sound classes without an architectural split. When frozen, it achieves a lower SELD score on real recordings than supervised baselines and improves spatial audio captioning and retrieval, with the largest gains in spatial accuracy. These results demonstrate a path toward robust, general-purpose spatial audio representations that jointly encode what is present and where it occurs in space.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.