EnvTTS: Towards Independent Speaker and Environment Cloning in Text-to-Speech
Abstract
Zero-shot text-to-speech (TTS) has made substantial progress in cloning unseen speakers from reference speech. However, the reference speech also captures acoustic characteristics of its recording environment, making speaker identity and environment difficult to control independently. Environment-controllable TTS (EnvTTS) uses separate speaker and environment references to synthesize a target speaker in a specified acoustic environment. We characterize the acoustic environment by two components: additive background sound and acoustic effects associated with reverberation. Evaluating EnvTTS requires assessing both speaker preservation and environment reproduction, yet systematic evaluation remains challenging due to the lack of a dedicated benchmark and metrics that separately characterize background and reverberation similarity. To address these gaps, we introduce EnvTTS-Bench, a benchmark of 3,000 evaluation instances with independently specified speaker and environment references drawn from real recordings, and EnvSim, an environment-similarity metric that separately measures background and reverberation similarity. We also establish two baselines: a training-free cascaded approach and EnvTTS-DiT, a flow-matching TTS model conditioned on separate speaker and environment references. Objective and human evaluations show that EnvSim aligns strongly with human judgments for separately measuring background and reverberation similarity. System evaluations on EnvTTS-Bench further show that the cascaded approach provides a practical baseline for EnvTTS, while EnvTTS-DiT achieves stronger overall performance; together, they provide useful reference systems for future EnvTTS research.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.