Towards End-to-End Conversational Multi-Talker Speech-to-Speech Translation
Abstract
We present COMUS, an end-to-end conversational multi-talker S2ST framework that leverages full contextual information to jointly model semantic meaning, speaker attribution, and prosodic characteristics for end-to-end multi-talker S2ST. Built on Qwen3-Omni, COMUS adapts the Thinker to autoregressively predict a single stream of text, speaker-label, and BiCodec speech tokens, replacing the original multi-token speech-generation module. Specifically, given the entire recording, COMUS predicts speaker profile tokens for each speaker alongside semantic tokens tagged with speaker labels, indicating speaker turns within a single autoregressive generation process. To provide the required supervision, we construct SODA-SoulX, comprising approximately 26K hours of paired speaker-consistent conversational speech across English and Chinese. Experiments on real conversational datasets, including AMI and AliMeeting, demonstrate that COMUS achieves speaker-attributed ASR-BLEU of 31.9 and 20.4 for English-to-Chinese and Chinese-to-English translation, respectively, outperforming all baseline methods. Further evaluations show translation advantages under noise and overlapping speech. Subjective listening tests show consistent advantages of COMUS in speaker similarity, translation accuracy and conversational naturalness over baseline methods.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.