Unified Target-Speaker ASR with Text and Enrollment Speech Cues
Abstract
Target-speaker automatic speech recognition (TS-ASR) aims to recognize speech from a designated speaker while suppressing interfering speakers in multi-talker environments. Conventional TS-ASR systems typically rely on an enrollment utterance to specify the target speaker, requiring additional speech from the same speaker at inference time. Text-guided approaches provide an alternative by exploiting known lexical content, such as a wake word, to locate the target speaker directly from the observed speech. However, these two forms of target guidance are typically studied independently, despite providing complementary information about the target speaker. In this work, we propose a Unified Dual-Cue TS-ASR framework that accommodates text cues, enrollment speech, or their combination within a single model. The text cue interacts with the mixture representation to extract target-speaker information conditioned on known lexical content, while an independent enrollment utterance provides complementary speaker information from a different utterance of the same speaker. Cross-attention cue-conditioning modules are embedded inside the shared Conformer blocks to condition ASR on the designated speaker and suppress interfering speech. The shared conditioning interface supports inference with text cues, enrollment speech, or their combination. During dual-cue training, both modalities are supplied and negative-cue sampling provides cue-validity supervision. Experiments on 30,000 two-speaker evaluation mixtures cover five recording and domain conditions and four oracle text-cue lengths. With five-character text cues, the proposed concatenated dual-cue method achieves an overall character error rate (CER) of 8.80%, compared with 17.32% and 29.06% under text-only and enrollment-speech-only inference, respectively. It also outperforms parallel dual-cue fusion, which obtains 9.49% CER, and yields lower dual-cue CER across all five evaluation subsets. These results demonstrate the advantage of jointly exploiting complementary lexical and speaker information for target-speaker ASR.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.