acceptodds
Under review as a conference paper at ICLR 2027

DISENTANGLING AUTHORED SPACE IN 7.1.4 MUSIC: A LARGE-SCALE CORPUS AND A NEURAL SPATIAL FILTERING CODEC

Abstract

Spatial intelligence has so far dealt with space that has a physical referent; the space people hear in music is authored, since every channel assignment in a 7.1.4 master is a mixing decision. Learning it has lacked a corpus, a representation that holds the space apart from the content, and a ground truth to check one against; to our knowledge, the largest channel-based 7.1.4 music dataset available for research has ten tracks. We produce 125,557 tracks and 8,392 hours of 7.1.4 music from Creative Commons–licensed stereo with an upmixing model trained on human masters, which is closer to them than a commercial upmixer on all five objective metrics and is rated by listeners between the two. Since a mix shapes its content through the relations between channels, the space can be carried as the filtering that turns a mono reference into twelve channels: our neural spatial filtering codec (NSFC) leaves the reference to a pretrained mono codec and transmits the space at a nominal 6 kbps (13.75 kbps in total), optionally with a flow-matching postfilter. Without ground truth, we test the separation by its behaviour, in downstream tasks built from the corpus that fix the reference and change only the spatial codes, so that they apply to any representation learned on it that carries the space apart from the content. Within the domain of the corpus, the stream carries spatial structure beyond static per-channel gains, transfers across musical content and, when substituted, moves the mix in the intended direction. Authored space thus becomes something that can be learned, separated from content and tested.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.