Emotional Enhancement of Cloned Voices via Speaker-Embedding Editing
Abstract
Native emotion instructions guide speech planning, but do not directly specify how the acoustic condition derived from a reference voice should change. We introduce a lightweight, reference-conditioned speaker-embedding editor for enhancing emotional expression in frozen text-to-speech synthesis. Trained on paired synthetic embeddings, the editor predicts a residual from the reference embedding and requested emotion, allowing both direction and magnitude to adapt to the reference rather than applying a shared shift. A scalar controls editing strength, and the edited condition enters CosyVoice 3 through its native acoustic interface without updating the synthesizer or requiring target-emotion recordings at inference. By holding speech tokens, acoustic prompts and decoding noise fixed, the method enables acoustic realization to be adjusted independently of changes to the speech plan. Experiments across Chinese and English, in both zero-shot and native emotional-instruction settings, show approximately 6.3–18.2% relative gains in mean target-emotion cosine. Bilingual listening evaluation further supports improved emotion appropriateness and conditional naturalness non-inferiority, alongside reduced speaker similarity at higher editing strength. These findings demonstrate a complementary approach to emotional enhancement that combines reference-dependent editing with adjustable control through a pretrained acoustic interface.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.