acceptodds
Under review as a conference paper at ICLR 2027

DeFormer: Latent-based music source separation beyond four stems.

Abstract

In music source separation, we aim to decompose a stereo music recording into a set of individual audio tracks, each isolating a specific instrument or group of instruments. Currently, most systems separate a given mixture considering a four-stem convention: drums, bass, vocals and a residual other that absorbs every other instrument. In this work, we introduce DeFormer, a music source separation model capable of separating a given mixture to 12 stems, decomposing at the latent representation level. Expanding the task from 4 to 12 demands greater sensitivity to instrument-specific cues due to finer distinctions in the target distribution, e.g., sparsity, spectral (or timbral) entanglement, etc. DeFormer addresses these challenges by combining a novel latent-separation architecture with a training recipe adapted to sparse and skewed supervision. On our held-out set, DeFormer reaches 6.59 dB mean cSDR over 12 targets, 1.06 dB ahead of the closest of four baselines we retrain at the same target count. DeFormer also demonstrates a notable gap in subjective performance on both general audio quality and adjacent source bleeding, measured in a conducted human study involving domain-expert annotators with professional music production backgrounds. Code and model checkpoints will be released upon acceptance. Samples page: https://deformerpaper-dev.github.io/DeFormer

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.