ViMed-Struct: Structured EMR Generation from Vietnamese Clinical Speech
Abstract
Clinical documentation requires physicians to repeatedly convert information conveyed during patient encounters into structured electronic medical record (EMR) fields, creating a substantial workload. Automating this process from speech is promising, yet progress is limited by the scarcity of paired clinical audio and structured EMR annotations. This bottleneck stems not only from privacy and practical constraints on clinical audio recording, but also from the high cost of structured data annotation, particularly for low-resource languages. To address this gap, we introduce a framework for constructing aligned clinical audio - structured EMR datasets from hospital records. By deriving structured targets directly from existing EMR data and separating them from speech realization, the framework removes the need for costly manual structured annotation while enabling scalable diversification of speakers, speaking styles, and acoustic conditions. Using this framework, we build ViMed-Struct, to the best of our knowledge the first large-scale Vietnamese speech-to-structured-EMR dataset, comprising 2,572 clinical cases and 2,892 recordings totaling 47.99 hours, collected from 58 speakers under matched quiet and noisy conditions. We further establish a benchmark with a field-aware evaluation metric that measures both correct field selection and extracted value quality, providing a more faithful evaluation than conventional text-similarity metrics alone. Finally, we propose an ASR - Information Extraction framework for speech-to-structured-EMR prediction and compare it with representative existing approaches. Experiments on ViMed-Struct show that our method achieves up to 94% macro-averaged field-presence F1, demonstrating strong performance in structured extraction across diverse backbone models.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.