acceptodds
Under review as a conference paper at ICLR 2027

Towards End-to-End Conversational Multi-Talker Speech-to-Speech Translation

Abstract

We present COMUS, an end-to-end conversational multi-talker S2ST framework that leverages full contextual information to jointly model semantic meaning, speaker attribution, and prosodic characteristics for end-to-end multi-talker S2ST. Built on Qwen3-Omni, COMUS adapts the Thinker to autoregressively predict a single stream of text, speaker-label, and BiCodec speech tokens, replacing the original multi-token speech-generation module. Specifically, given the entire recording, COMUS predicts speaker profile tokens for each speaker alongside semantic tokens tagged with speaker labels, indicating speaker turns within a single autoregressive generation process. To provide the required supervision, we construct SODA-SoulX, comprising approximately 26K hours of paired speaker-consistent conversational speech across English and Chinese. Experiments on real conversational datasets, including AMI and AliMeeting, demonstrate that COMUS achieves speaker-attributed ASR-BLEU of 31.9 and 20.4 for English-to-Chinese and Chinese-to-English translation, respectively, outperforming all baseline methods. Further evaluations show translation advantages under noise and overlapping speech. Subjective listening tests show consistent advantages of COMUS in speaker similarity, translation accuracy and conversational naturalness over baseline methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.