acceptodds
Under review as a conference paper at ICLR 2027

MRareBench: A Multimodal Rare-Disease Benchmark for Evidence-Diagnosis Correspondence

Abstract

Multimodal large language models (MLLMs) now approach specialist-level performance on chest radiograph interpretation, dermatology and pathology image classification, and medical visual question answering. These gains, however, come from common diseases, where the clinical narrative alone often narrows the differential. Rare diseases reverse these conditions. Each disease is individually uncommon and sparsely documented, prevalence gives little guidance across thousands of candidates, and the decisive findings are frequently visual and distributed across several images and modalities. Rare disease therefore offers a stringent test of whether a MLLM reaches a diagnosis from evidence or from prior familiarity. Current rare disease benchmarks describe patients in text and omit medical images, while multimodal medical benchmarks supply images but concentrate on common conditions. To address this limitation, we introduce MRareBench, the first benchmark for multimodal rare disease diagnosis. MRareBench comprises 908 evaluation items spanning multiple medical images and diverse imaging modalities, and pairs forward diagnosis with evidence verification under controlled disclosure of the diagnosis. Evaluation of 37 closed-source, open-source, and medical MLLMs shows that visual evidence improves rare disease diagnosis. A gap remains between evidence recognition and open diagnosis. MLLMs often identify relevant visual evidence after receiving the rare disease diagnosis but are less reliable in using multimodal evidence to identify the disease independently. The results indicate a difficulty in inferring rare disease diagnoses from multimodal evidence. MRareBench supplies the measurement needed to advance multimodal diagnosis in rare-disease settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.