Silent Tags: Expressive Speech without Language-Backbone Updates
Abstract
With the rapid development of omni large language models, users can converse directly with assistants through speech. However, most existing models render every response in the same flat prosody, regardless of whether the user is distressed, excited, or amused, limiting the user experience in human-machine interaction. To inject expressiveness, the prevailing approach trains the language backbone, namely the Thinker, to jointly predict paralinguistic tags alongside text. We find that this consistently degrades reasoning on benchmarks such as MMLU and VoiceBench. A key challenge therefore lies in how to make a speech assistant expressive without compromising the reasoning ability of its language backbone. To address this challenge, we propose Silent-Tags, a Thinker-agnostic framework that leaves the Thinker entirely frozen and trains only the Talker. Our key insight is that a well-trained omni Thinker already encodes, in its hidden states, contextual signals required for emotional and paralinguistic inference. Silent-Tags equips the Talker with lightweight cross-attention layers that selectively tap these hidden states for rich emotional and paralinguistic cues, supervised by a tag-controllable TTS teacher together with a segment-level auxiliary classifier. Silent-Tags introduces no extra Thinker decoding steps and no tag-token generation at inference time; the only added overhead is the Talker's cross-attention into already-cached Thinker states. On ParaS2SBench and EmergentTTS-Eval, Silent-Tags consistently outperforms explicit-paradigm baselines on expressiveness. Audio samples, code, and the full tag vocabulary are available at [link](https://anonymous-3911.pages.dev/).
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.