MedSonor: Learning to Use Clinical Audio in Medical LLMs
Abstract
Clinical audio conveys medical information that transcripts may not preserve, but using it in medical responses requires interpreting it alongside clinical requests and dialogue context. Existing medical audio supervision largely targets classification or recognition, while general audio-language models lack specialized medical knowledge and direct feedback for open-ended responses. We present MedSonor, a medical dialogue model integrating medical knowledge, spoken dialogue, and clinical audio understanding. We pair authentic medical audio with clinical requests and combine these audio-integrated dialogues with collected medical audio and synthesized spoken dialogues for supervised fine-tuning, followed by two stages of rubric-guided reinforcement learning. We introduce MedVoice-Dial and MedVoiceAID to evaluate responses within supplied multi-turn medical contexts and to requests involving clinical audio, respectively. Across seven medical benchmarks, MedSonor outperforms the evaluated general-purpose open-weight audio-language baselines and ranks highest in aggregate physician preference evaluation, while broadly retaining general speech capabilities.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.