AudioFM: Towards General Purpose Audio Representation Learning
Abstract
Learning audio representations that effectively bridge holistic global soundscapes with fine-grained acoustic moments is a long-standing goal in self-supervised learning. Recording-level contrastive objectives encourage agreement between pooled audio representations, but those objectives do not explicitly supervise time-frequency local representations. We introduce AudioFM, a self-supervised framework that combines recording-level contrast with patch-wise masked latent discrimination. Using visible acoustic context, a student predicts representations of masked patches and identifies their corresponding full-view teacher targets among candidates from other recordings. We further incorporate weaker supervision from eligible temporal neighbors to encourage local temporal consistency. Extensive evaluations show that AudioFM achieves state-of-the-art performance across diverse audio domains in linear probing, fine-tuning, and zero-shot transfer, demonstrating superior event detection, robust scene-level understanding, and strong domain generalizability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.