MultiJEPA: Learning Music Representations Across Temporal Resolutions
Abstract
We present MultiJEPA, a multi-resolution joint-embedding predictive architecture for learning representations of signals whose semantics occur at multiple scales. Rather than committing to a single resolution, MultiJEPA trains one encoder across dynamically aggregated resolutions. In music, semantics span timescales from the millisecond-scale precision required for onset and beat localization to harmonic context and musical form unfolding over tens of seconds to minutes. Music is therefore an attractive domain for studying how temporal resolution relates to learned semantics. Specifically, a controlled five-resolution analysis reveals a strong context–detail tradeoff: coarse representations favor tonal and structural understanding, while fine representations favor beat tracking and transcription, with no single resolution dominating. MultiJEPA recovers this fixed-resolution frontier with a single encoder queried at four temporal rates. Simple four-query fusion exploits complementary information across resolutions, outperforming the compared audio-only SSL encoders on 15 of 16 probing metrics spanning nine downstream tasks. We further propose a short fine-to-coarse self-alignment phase that organizes these views into a nested representation and improves cross-resolution probe transfer. Together, these results establish resolution as an explicit design axis for music understanding models, with MultiJEPA as a concrete method to extend the pareto frontier of the context-detail tradeoff.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.