MedEviBench: Evaluating Large Language Models for Evidence-Based Medicine
Abstract
Medical question answering benchmarks measure what language models know, but provide a limited view of how well they use evidence to answer questions from clinical practice. Evaluating this ability requires questions that reflect physicians' information needs and measures that assess the evidence supporting a clinical answer. We introduce MedEviBench, a Chinese-language benchmark of 240 physician questions spanning 12 medical specialties and ten clinical intents. Given a clinical question and reference passages, a model must select relevant evidence and produce an answer with citations supporting its conclusions. We evaluate these answers using nine criteria grouped into evidence quality, clinical content, and answer presentation. Blinded physician comparisons on questions sampled from the benchmark yield a Spearman correlation of 0.80 between automatic and physician system rankings. We evaluate 18 LLMs, including 13 general-purpose and five medical models. Our results show that producing reliable clinical answers remains difficult even when reference evidence is provided. Among the general-purpose models, clinical content scores range from 47.31 to 82.96 out of 100, falling below evidence quality for 12 of the 13 models and presentation for all models. These differences are not consistently reflected in medical knowledge scores: an exploratory comparison across 12 models finds a moderate correlation between AA-Omniscience Health and clinical content, but almost no linear correlation with evidence quality. To examine where models fail, we analyze judge explanations for zero-scored criteria and find recurring errors in clinical conclusions and inferences that exceed the available evidence. MedEviBench tests whether models can turn reference evidence into sound clinical answers, making these failures part of the assessment of medical language models.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.