acceptodds
Under review as a conference paper at ICLR 2027

CE-TSE: Cross-Emotion and Physical Acoustics for Robust Target Speaker Extraction

Abstract

Extracting the target speaker's voice from noisy mixtures environment is a core challenge for intelligent auditory systems. Target Speaker Extraction (TSE) aims to mimic the human "Cocktail Party Effect". When humans extract target speech in noisy environments, they rely on not only speaker identity, but also spatial acoustic cues and emotional information. However, existing TSE benchmark datasets lack both realistic acoustic conditions such as room reverberation and spatial geometry, and complete emotional annotations. Therefore, we manually annotate emotional labels for over 2,300 utterances from the REAL‑TSE dataset, consuming approximately 58 person‑hours of annotation effort. Then we incorporate real room impulse responses and noise to perform physical spatial simulation to build the Cross-Emotion Target Speaker Extraction dataset (CE-TSE Dataset). Building upon this dataset, we propose the Cross-Emotion Target Speaker Extraction Network (CETSE-Net). CETSE-Net is equipped with an emotion disentanglement module (CE), which decouples emotional features from speaker identity representations. Experimental results show that CETSE-Net achieves consistent improvements across all evaluation metrics: +1.844 dB on SI-SNR, +1.736 dB on SDR and +0.42 on PESQ.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.