SEPADAPTER: A SEPARATION ADAPTER FOR STREAMING AND OFFLINE MULTI-TALKER ASR BACKBONES
Abstract
Multi-talker speech recognition must transcribe a target speaker’s speech while limiting interference from overlapping voices. We propose SEPADAPTER, an activity-conditioned residual module placed after convolutional subsampling and before the automatic speech recognition (ASR) encoder. It uses a three-channel, per-frame target-relative activity signal indicating whether the target, a competing speaker, and remaining competitors are active, together with temporal context, to correct acoustic features through a residual connection. We implement the core design in FastConformer, Zipformer, and Whisper, with separately trained weights and backbone-specific context handling for streaming and offline recognition. In a parameter-matched FastConformer comparison, the 3.78M-parameter dual-path module reduces concatenated minimum-permutation word error rate (cpWER) by 0.356 percentage points (14.52% to 14.17%, 2.4% relative) over framewise conditioning. This reduction appears in all twelve seed-by-lookahead cells across three paired training runs. Diagnostics on one paired training run show gains concentrated in overlapped words and increased errors when adapter history is blocked. Insertion after encoder block 8 raises cpWER by about 7 points under the shared recipe, even though the advantage over framewise conditioning grows. The separately optimized FastConformer system achieves lower cpWER than published streaming self-speaker adaptation (SSA) results in all twelve LibriSpeechMix conditions, by 0.22–5.47 points. The mean gap grows from 0.66 points at 1120 ms ASR lookahead to 4.16 points at 80 ms, although training and diarization recipes differ. Our Whisper system, which also uses self-enrollment and encoder finetuning, improves over locally re-evaluated SE-DiCoW on three LibriMix conditions under oracle activity. With DiariZen activity, it reduces LibriSpeechMix three-speaker time-constrained cpWER (tcpWER) from the published 21.1% to 19.84%; results vary across domains. The controlled results support temporal activity-conditioned correction as a complement to the encoder’s temporal modeling.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.