acceptodds
Under review as a conference paper at ICLR 2027

HEAR: Human-likeness Evaluation with Acoustic Reasoning for Generative Speech

Abstract

Assessing the perceived human-likeness of generated speech is important but challenging. Listening tests are costly to scale, quality-oriented metrics do not directly measure this target, and raw-audio AudioLLM judges offer limited explicit acoustic evidence. We propose HEAR, a training-free method for Human-likeness Evaluation with Acoustic Reasoning. HEAR uses off-the-shelf tools and explicit rules to convert acoustic measurements into semantic tags organized at global, word, and segment levels. This representation links measured behavior to its location and surrounding speech. A dedicated human-likeness prompt guides an AudioLLM to judge the original audio using the target text and evidence, producing an overall human-likeness score, eight dimension-specific scores, and evidence-grounded reasons. On Humanlike-100, HEAR raises Qwen3-Omni-Thinking's Pearson and Spearman correlations with mean human overall ratings to and those of matched audio-only judging, respectively. HEAR improves overall agreement with human ratings across all five evaluated AudioLLM backbones. Ablation studies support the contributions of structured acoustic evidence and the human-likeness evaluation rubric. Beyond human-likeness evaluation, task-adapted HEAR improves audio-only judging on speech preference and perceptual-quality benchmarks. Code is available at https://github.com/HEAR-6/HEAR.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.