ControlAnything: Unified word-level, Intensity-Aware, and Composable Control for Expressive Text-to-Speech
Abstract
Emotional expression in human speech varies across words, exhibits varying intensity levels, and often blends multiple affective states. Reproducing this flexibility in text-to-speech (TTS) requires joint control over where emotions occur, how strongly they are expressed, and how they are combined, while applications such as voice assistants further demand fast generation for responsive interaction. However, existing methods typically address these control dimensions separately, and fine-grained word-level control has been explored primarily in autoregressive (AR) TTS, leaving unified control in efficient non-autoregressive (NAR) architectures largely underexplored. We present ControlAnything, a unified framework for word-level, intensity-aware, and composable emotional control that preserves NAR inference efficiency. First, Soft-boundary Anchor Intervention (SAI) maps target text spans to acoustic regions and progressively reinforces temporal anchoring as acoustic representations evolve during generation. Second, Emotion Concept Modulation (ECM) learns distinct emotion-specific transformations within a shared modulation mechanism, enabling fine-grained intensity adjustment and flexible multi-emotion composition. Together, SAI and ECM support localized intensity adjustment and emotion composition across different pretrained TTS models. Extensive experiments demonstrate state-of-the-art controllability in word-level localization, emotion-intensity adjustment, and multi-emotion composition, while preserving speech quality and generation efficiency comparable to the baselines. Audio demonstrations are available at https://controlanything-review-demo.vercel.app/.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.