CoReSep: Input Composition for Unifying Speech Separation and Target Speaker Extraction with a Pretrained Speech Foundation Model
Abstract
Speech separation (SS) and target speaker extraction (TSE) share the same underlying goal of disentangling overlapping speakers, providing a natural basis for unified modeling. However, existing unified systems either extend acoustically driven separation backbones with speaker-selection mechanisms, tying source attribution primarily to acoustic evidence, or adopt single-stream language-model generation that requires repeated inference for multi-speaker SS without explicit cross-stream coordination. To address these limitations, we introduce CoReSep, a unified SS and TSE system built around two key ideas. First, separation is performed by a pretrained speech foundation model (SFM) itself, whose contextual prior provides high-level constraints on source attribution beyond acoustic evidence alone. Second, a novel input composition mechanism enables joint multi-stream modeling by specifying each task entirely through the input: replicating the mixture representation instantiates multiple jointly decoded streams for SS, whereas prefixing an enrollment representation anchors the single output stream to the target speaker for TSE. In this way, a standard single-stream SFM serves as the unified separator without architectural modification. A conditional flow-matching module then restores acoustic detail by conditioning on both the separated representation and the mixture Mel spectrogram, after which a vocoder reconstructs the waveform. Experimental results demonstrate that CoReSep outperforms previous task-specific and unified systems in linguistic integrity on both SS and TSE, while maintaining competitive perceptual quality. On Libri2Mix, it achieves differential word error rates (dWERs) of 0.85%/3.65% for SS and 1.22%/3.93% for TSE on the Mix-Clean/Mix-Both subsets, respectively. It also maintains its advantage over a strong discriminative SS baseline on three- and four-speaker mixtures and outperforms it on all metrics on EchoSet, whose reverberation and random overlap ratios are closer to real-world conditions. Audio samples are available online.samehttps://anonymous.4open.science/w/CoReSep-Demo-4E0B/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.