Recognition-Aware Selection of Enhancement-Derived Transcripts
Abstract
Different speech enhancers supply a frozen recognizer with complementary corrections and conflicting errors. CARE-ASR selects among the resulting transcripts by combining expert risk estimates with transcript agreement and acoustic support. It learns word-error differences and calibrates replacement against a complete default policy. We evaluate four English corpora, then apply the original meeting-speech policy without recalibration to 4,096 project-held-out AMI recordings from 16 meeting–speaker groups. Word error rate falls from 30.040% for reverse ROVER to 27.860%, a 2.180-percentage-point reduction (95% group-bootstrap interval [1.455, 2.919]; adjusted ). Fifteen groups improve, with more recovered words, fewer erasures and fewer insertions overall. Same-cohort controls examine the representation: expert inputs help under both pairwise and MSE training, whereas a matched nonlinear selector using only eight expert risks makes 210 more errors. When expert predictions for the AMI fitting sources are generated out of fold, the gain over reverse ROVER persists, and masking the explicit expert inputs adds 240 errors. The findings support learning from expert judgements together with candidate evidence to exploit enhancement diversity without updating the recognizer.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.