EmoScale: Localized Emotion Intensity Control for Speech Synthesis via Twin-Forward Flow Matching
Abstract
Recent flow-matching and diffusion-transformer (DiT)-based text-to-speech (TTS) models achieve high naturalness, yet precise and localized control of emotional expression remains challenging. Existing instruction- or encoder-conditioned methods inject emotion through labels, embeddings, or prompts, while training-free steering methods extract emotion directions post hoc without explicitly enforcing disentanglement. These approaches provide limited control over the intensity and temporal extent of emotional expression, as the realised emotion can be influenced by the model’s internal representations and interactions between emotion conditioning and linguistic content. Moreover, conventional categorical emotion representations and natural-language prompts provide limited means to precisely specify how strongly and when an emotion should be expressed. To address these challenges, we propose EmoScale, a training paradigm that explicitly disentangles the velocity field of DiT-based flow-matching model into semantic and emotion components without requiring an external emotion embedding. First, we introduce a disentanglement regularizer that encourages the emotion velocity to be orthogonal to the semantic velocity. We further show that the resulting emotion representation enables continuous control of emotional intensity through linear interpolation of its corresponding acoustic attributes, including F0, periodicity and envelope, followed by speech reconstruction. This allows us to adjust the intensity of an emotional speech while preserving perceptual naturalness. We evaluate EmoScale against encoder-conditioned and training-free steering baselines in terms of naturalness, emotion recognition accuracy, content preservation, and localized emotion control. Results demonstrate that explicitly disentangling the velocity field enables precise and interpretable control over emotional intensity without requiring additional trainable auxiliary modules. The resulting RA scores exhibit a clear and consistent trend as the emotion intensity is varied, confirming the effectiveness of the proposed control mechanism.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.