Selective Context Utilization in Speech-LLM ASR: Lexical Verification and Semantic Disambiguation with GRPO
Abstract
Context can improve entity recognition in speech-enabled large language models (Speech-LLMs), but misleading context can also bias transcription toward unsupported words. Context additionally provides semantic evidence for written distinctions that pronunciation alone cannot resolve. We study these complementary demands through contextual entity recognition and Mandarin third-person pronoun disambiguation. We develop a three-stage framework combining general ASR fine-tuning, robust context fine-tuning with mismatched contexts, and task-specific group relative policy optimization (GRPO). The lexical objective uses phonetic hard negatives and quality-gated distractor rejection, while the semantic objective rewards correct ordered pronoun sequences alongside full-transcription accuracy. We construct a diagnostic benchmark suite with 1,116 lexical utterances and 617 pronoun-disambiguation utterances, supplemented by paired clean/misleading contexts and controlled semantic-context ablations. Our model achieves 97.76% entity recall and 89.70% pronoun macro-F1, exceeding the strongest evaluated pronoun baseline by 7.12 percentage points. Under misleading lexical context, the full pipeline improves entity recall while matching the backbone’s distractor false-positive rate. Semantic ablations show that the disambiguation gains depend on matched discourse, although general-ASR performance degrades unevenly across domains. We will release the benchmark upon acceptance. The code is available at https://anonymous.4open.science/r/SCU-ASR-8E2D.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.