Voice Acting Arena: Evaluating Expressive Speech Across Acting Challenges
Abstract
Expressive text-to-speech models increasingly follow instructions, yet authentic voice acting remains difficult to compare: saying the right words is not the same as performing a scene. We introduce the Voice Acting Arena, a human-preference evaluation for complex acting directions. Its challenges span 40 fine-grained emotion targets, selected vocal-delivery directions, and multi-step scenes across ten acting domains. A language-model director translates each challenge into the controls supported by a particular speech system; listeners then compare anonymized takes on overall acting, instruction fit, authenticity, and, when requested, a vocal reaction. In a first round of 560 listener votes on eleven model entries, Gemini 3.1 Flash TTS leads on every question. Our Humaneness Voice baseline, released with commercially usable open weights, sits in the middle: with ten-candidate selection its rating cannot be separated from OpenAI GPT Audio 1.5, DramaBox, or Cartesia Sonic 3.5. To curate its training data and rank candidate takes, we train a Genuineness Score for spontaneous, authentic-sounding speech and a Vocal Burst Blending Score for how naturally laughs, sighs, and similar bursts fit the surrounding speech. A first listener check supports the genuineness score only weakly and does not yet support the blending score, so neither is treated as a measure of acting. We release the challenge design, evaluation protocol, bilingual data views, and baseline components, and keep live descriptive results separate from a frozen benchmark.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.