AdaSpeechRAG: Auditing Evidence for Supplied-Context ASR
Abstract
Automatic speech recognition (ASR) systems must sometimes decide whether to use contextual hints or a more expensive recognizer. Applying the strongest recognizer to every utterance wastes computation, and indiscriminate context can insert unsupported words. Prior work usually evaluates these actions separately, which makes their quality, latency, and risk trade-offs difficult to compare. We introduce AdaSpeechRAG, an auditable interface for a supplied-context decision: keep a cheap transcript or add a supplied entity prompt. On the 3,750-utterance held-out partition of ContextASR-Bench, fixed prompts reduce word error rate (WER) from 26.66% to 22.11% and improve entity-level F1 from 66.15% to 83.98%. A rule chosen on the training and development partitions reaches 21.94% WER while prompting 60.27% of test utterances. Relative to fixed prompting, its paired absolute character error rate (CER) difference is -0.37%; prompt use and recorded p95 real-time factor (RTF) also fall in one execution setting. Its WER difference remains unresolved, while entity-level F1 falls to 82.96%. These matched comparisons expose a quality, entity, and prompt-use trade-off that a WER-only summary would miss. AdaSpeechRAG records the scope and provenance of each action so that supplied-context results and stored action counts remain distinguishable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.