CoCiTTS: Towards Character-Level Speaker Consistency in Zero-Shot TTS for Cross-Lingual Cinematic Dubbing
Abstract
Cross-Lingual Cinematic Dubbing aims to synthesize target-language speech from source-language speech in movies or TV shows, which requires a character to retain a consistent voice across scenes with different emotional expressions. However, existing zero-shot text-to-speech (TTS) systems mainly optimize utterance-level speaker similarity and overlook character-level speaker consistency across multiple utterances. This limitation becomes particularly challenging in cinematic scenarios, where the same character frequently exhibits diverse emotional states, and emotion-rich reference utterances may lead to inconsistent speaker timbre across generated utterances. In this paper, we identify emotion leakage within speaker representations as an important factor contributing to speaker timbre drift. To address this issue, we propose CoCiTTS (Consistency-Oriented Cinematic Text-to-Speech), a zero-shot TTS framework that learns consistency-oriented speaker conditioning for cross-lingual cinematic dubbing. CoCiTTS introduces a Speaker Timbre Purification (STP) module based on adversarial disentanglement to suppress emotion leakage while preserving speaker identity. Furthermore, a Fine-grained Emotion Compensation (FEC) module models global and local emotional characteristics through an independent pathway, enabling expressive generation without contaminating speaker representations. In addition, we construct CineSpeech, an emotion-rich cinematic speech benchmark containing 136 characters, and establish consistency-aware evaluation methods for character-level timbre preservation. Extensive experiments on multilingual cross-lingual tasks demonstrate that CoCiTTS effectively improves speaker consistency while maintaining competitive speech quality, speaker similarity, and emotional expressiveness compared with existing zero-shot TTS systems.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.