acceptodds
Under review as a conference paper at ICLR 2027

X-Emo: Disentangled Speech Tokens for Cross-Lingual Zero-Shot Expressive Text-to-Speech

Abstract

Cross-lingual zero-shot expressive text-to-speech (TTS) aims to generate speech in a target language while preserving the speaker identity and emotion of a prompt utterance in another language. A key challenge is to transfer these attributes without entangling them with linguistic content and source-language characteristics. We propose X-EMO, a cross-lingual expressive TTS framework built on a disentangled discrete speech tokenizer. During tokenizer training, the decoder receives an additional utterance that shares the speaker and emotion of the input but contains different linguistic content. This makes speaker and emotional information directly available to the decoder, encouraging the discrete representation to primarily encode information required for linguistic reconstruction. The resulting tokens are modeled by an autoregressive language model, while a flow-matching acoustic model synthesizes speech conditioned on speaker and emotion embeddings from the prompt utterance. With approximately 65 hours of expressive speech added to bilingual read-speech data, X-EMO achieves competitive intelligibility, speaker similarity, naturalness, and emotion preservation across multiple languages, while maintaining stable cross-lingual performance across diverse prompt languages.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.