Unifying Medical Vision-Language Learning Across Imaging Modalities in Vietnamese: A Multimodal Dataset and Benchmark
Abstract
Medical report generation has advanced rapidly with Vision-Language Models (VLMs), yet progress remains constrained by the scarcity of high-quality paired medical image-report data, particularly for low-resource languages. Existing datasets are largely modality-specific, covering common modalities such as X-ray, CT, or MRI individually rather than jointly, while PET/CT remains underrepresented. Moreover, many datasets are assembled from figure screenshots or re-captured images rather than the native DICOM studies acquired in routine clinical practice, limiting their clinical fidelity and practical value. To address these limitations, we introduce **ViMed-Rad**, a large-scale multimodal radiology dataset jointly covering four major imaging modalities: X-ray, CT, MRI, and PET/CT. ViMed-Rad comprises 93,781 clinical studies paired with Vietnamese reports authored by radiologists, including 62,148 X-ray, 13,247 CT, 12,010 MRI, and 6,376 whole-body PET/CT studies. All studies are native DICOM acquisitions exported from hospital PACS archives, retaining acquisition metadata and spanning diverse anatomical regions and imaging protocols. Building on ViMed-Rad, we establish **ViMed-Bench**, a unified benchmark that couples clinical report generation and report-grounded VQA tasks with a multimodal framework integrating modality-specific visual representations into a shared language generation space, so that a single model serves all four modalities. Benchmarking existing medical VLMs reveals substantial performance gaps across imaging types, particularly for underrepresented modalities, together with a consistent degradation when reports are generated in Vietnamese. Experiments on ViMed-Rad show that the proposed framework accommodates heterogeneous planar, volumetric, and functional-anatomical imaging within one model, outperforming baselines across all four modalities while retaining most of the accuracy of modality-specific counterparts, and providing a scalable foundation for multi-modality medical VLMs rather than separate models for each imaging domain.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.