acceptodds
Under review as a conference paper at ICLR 2027

ASR Evaluation with Phonetic Embeddings

Abstract

Evaluating a speech system using mean or median Word Error Rate (WER) can obscure meaningful differences in performance among subgroups within the evaluation dataset, and among groups of speakers who will interact with the system. Prior work has shown that these disparities can be significant, and pose a challenge to adoption. However, to measure fine-grained disparities, developers typically rely on datasets that are annotated with speaker demographics. Such annotations are costly, noisy, and not causally connected to speech behaviors that trigger recognition error. In this work, we propose a new data-driven technique for measuring error disparity. It is fully automatic, derived from meaningful speech behavior, and directly correlated with drivers of error. Phonetic embeddings are a data-driven representation of speech that can be used to cluster speakers and measure error disparity among clusters. Our technique addresses the gaps in current evaluation methods with a speaker representation that directly correlates with error, and a fully automatic method for disparity analysis. In an evaluation of an off-the-shelf commercial ASR system, we show that this method finds larger regression coefficients, performance gaps with similar or greater effect size, and better cluster quality, compared to demographic methods.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.