HealthCoreBench: Measuring Evaluation Artifacts for Reliable Medical LLM Evaluation
Abstract
Medical evaluation of large language models (LLMs) increasingly relies on heterogeneous benchmarks, yet benchmark failure is not always a reliable indicator of limited medical intelligence: unsolved samples may reflect genuine capability gaps or evaluation artifacts such as ambiguity, incorrect references, or outdated knowledge. We formulate medical benchmark construction as a **failure attribution** problem: identifying samples where failures can be meaningfully attributed to capability limitations. We introduce **HealthCoreBench**, a unified medical evaluation benchmark distilled from **47** language benchmark sources spanning **165,120** samples, of which only **13,987** (**8.47%**) are retained. HealthCoreBench uses progressive model-based filtering to identify frontier-challenging samples, followed by clinician validation to flag and correct unreliable failures. Auditing large-scale medical benchmark collections reveals that: 1) unreliable failures are prevalent, and difficulty filtering alone can even enrich evaluation artifacts, since selecting samples that strong models fail on preferentially retains defective rather than genuinely hard questions; 2) these artifacts create false successes and shift model scores unevenly, with physician validation alone changing overall scores by up to **5.18** points across models and distorting apparent medical capability; and 3) retaining only clinician-validated failures preserves model rankings while cutting average token consumption by **77.17%** (from **82.82M** to **18.91M** tokens). Overall, unreliable failures undermine evaluation validity, and failure attribution offers a path forward. We release HealthCoreBench to shift medical evaluation from “unsolved as failure” toward “failure attributable to capability limitations.”
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.