Search-ASR: Training Speech Agents to Search for Entity-Aware Speech Recognition
Abstract
Large language models have advanced automatic speech recognition (ASR), improving multilingual and dialectal recognition and contextual understanding. However, despite accurately transcribing most of an utterance, ASR systems still struggle with emerging and long-tail entities owing to static parametric knowledge and limited coverage of prebuilt hotword retrieval resources. This paper introduces Search-Augmented Agentic ASR (Search-ASR), which trains speech agents to autonomously query Web search engines using acoustic cues and utterance context, retrieve candidate names and semantic information, and combine this evidence with the audio to produce transcriptions. We further develop a multi-stage post-training framework and a multi-agent data synthesis system that establish search capabilities through supervised finetuning and reinforcement learning while reducing redundant searches and unnecessary rewriting. Experiments on English and Chinese entity-focused ASR benchmarks show that Search-ASR outperforms all evaluated open-source baselines and remains competitive with evaluated closed-source systems. Compared with its initialization model, Qwen3-Omni, Search-ASR reduces English word error rate and Chinese character error rate by 53.1% and 25.5%, while improving entity recall in both languages.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.