NumPipe: Automatic Generation, Vocalization and Evaluation of Spoken Numbers
Abstract
In the last year AI agents, meeting summarizers and similar assistants built on automatic speech recognition (ASR) have spread quickly. Their output is used directly and often unchecked, so a wrong number in a transcript is a serious error. Word error rate (WER), the fraction of words an ASR model gets wrong, reflects this problem poorly, because numbers are rare in text. In the standard test sets we count only about 13 numbers per 1000 words. Instead, we measured the pure number transcription accuracy. By this measure, two widely used open ASR models with low general error rates get only 80–88% of hard dictated numbers right. To measure and fix this at scale, we present NumPipe, an automatic pipeline with no human in the loop. It generates numbers with known values and vocalizes them with preset and zero-shot voices. Independent models then judge every clip, and the model under test is scored on the value it recognized. Because the same stages also produce training data, the pipeline can repair the weakness it measures. A 2000-step finetune on this data raises two models of different architectures from 80–88% to 98.8–99.8% on hard dictated numbers. On the public Numb3rs set they reach 97.6–98.8%. The price is 0.7 to 1.2 points of general WER. NumPipe is released as a set of NeMo Curator stages, and its test sets form a benchmark for numeric accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.