acceptodds
Under review as a conference paper at ICLR 2027

Dr.EMO: De-biased Reward Optimization for Emotionally Expressive Text-to-Speech

Abstract

Achieving natural and controllable emotional expression remains a key challenge in text-to-speech (TTS) synthesis. Reinforcement learning offers a promising approach through emotion reward optimization. However, such optimization relies on reward models that can reliably recognize emotions in synthesized speech. Since human-verified emotional labels for synthesized speech are scarce, existing methods typically employ speech emotion recognition (SER) models trained on real human speech as reward models. Yet, we identify a substantial misalignment between such models and human judgments on synthesized speech. To address this issue, we curate an open-source dataset of manually filtered synthesized speech with clear emotional expression and train a de-biased emotion reward model that better aligns with human judgments. Based on this model, we introduce Dr.EMO, a multi-objective group relative policy optimization framework that enhances emotional expressiveness while preserving general TTS capabilities. Experiments demonstrate that Dr.EMO achieves state-of-the-art performance in conveying target emotions, with over 90% of synthesized utterances exhibiting clear human-evaluated emotional expression. Our findings highlight the importance of reliable, de-biased rewards for effective emotional TTS training.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.