acceptodds
Under review as a conference paper at ICLR 2027

ExpressTTS-Bench: Benchmarking the Scenario-based Expressiveness of Text-to-Speech Models

Abstract

Recent Text-to-Speech (TTS) systems have achieved substantial progress in sound quality and integrity, yet their ability to generate natural and appropriate highly expressive audio based on a given scenario is still very limited. Meanwhile, the existing evaluation of expressiveness still relies on costly and inefficient human evaluation. Moreover, existing expressiveness evaluation typically overlooks the scenario and textual context that fundamentally shape how expressiveness should be perceived. To this end, we introduce ExpressTTS-Bench, a fully automated benchmark for evaluating the scenario-based expressiveness of TTS models. ExpressTTS-Bench constructs diverse test data through a hierarchical scenario tree covering five representative expression modes and four task formats, resulting in test samples across English and Chinese. To enable fine-grained and scalable evaluation, we decompose expressiveness into nine sub-dimensions spanning emotion, prosody, and paralanguage, and introduce a LALM-as-a-Judge framework with automatically generated, expression-mode-specific scoring rubrics. We evaluate six representative open-source and closed-source TTS models and reveal substantial differences in scenario-based expressiveness across models, languages, instruction formats, expression modes, and sub-dimensions. Consistency analyses and human evaluation further support the reliability of ExpressTTS-Bench. ExpressTTS-Bench provides a scalable framework for systematically evaluating the scenario-conditioned expressive capabilities of modern TTS systems.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.