acceptodds
Under review as a conference paper at ICLR 2027

CARE-Med: A Capability-Resolved Benchmark for Medical Perception across Photographic and Physiological-Signal Images

Abstract

Medical multimodal large language models (MLLMs) are still compared mostly by a single aggregate accuracy, which reveals that a model errs but not which competency it lacks. Existing benchmarks also concentrate on sensor-reconstructed imaging (radiology, pathology, ultrasound, and microscopy), leaving camera-captured images and physiological-signal visualizations under-tested. We present CARE-Med, a 2,023-question Chinese-language benchmark over five under-tested domains: body fluids, in-vitro diagnostic (IVD) reagents, endoscopy and surgical fields, ex-vivo organs and tissues, and physiological-signal visualizations. Three of them (macroscopic body fluids, IVD-reagent readouts, and gross ex-vivo organs and tissues) have, to our knowledge, no dedicated vision–language evaluation, and a self-built four-level image taxonomy quantifies this coverage bias across seven prior benchmarks. Every question is mapped to leaves of a 34-leaf capability tree whose top level takes three knowledge types (factual, conceptual, procedural) from Bloom’s revised taxonomy, and every question, answer, and label is verified by clinicians. Evaluating 14 frontier and medical MLLMs, we find that a single score, even when split by knowledge type, is misleading: within each frontier model the three knowledge types differ by only 1–3 points, yet its capability leaves span more than 30 (GPT-5.5 ranges from 53% to 89% and reaches only 71% overall). The three frontier models fail on largely the same capability leaves (ρ≈0.8), a shared weakness in concrete visual determination alongside strong explanation, and medical-specialized models redistribute competence across the tree rather than closing this gap. CARE-Med turns an opaque score into a diagnosis of where medical perception breaks.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.