acceptodds
Under review as a conference paper at ICLR 2027

MedStat-Bench: Evaluating Diverse Multimodal Medical Intelligence from Perception to Agentic Reasoning

Abstract

Multimodal large language models (MLLMs) are often evaluated on a single medical modality or task, leaving their behavior across clinical data types and workflows poorly characterized. We present MedStat-Bench, a broad stress test with 430 curated samples from 6 public datasets and 5 institutional cohorts covering about 10 clinical conditions and care settings. It spans clinical text, 2D images, volume-derived views from 3D studies, and sampled video, with tasks including single- and multi-label classification, structured clinical text generation, regression, and 2D/3D spatial localization. We release a self-contained public track with fixed evaluation materials; institutional cohorts test the same framework on locally collected clinical data. An initial agentic track contains 29 scenarios across 11 patient cases requiring evidence retrieval, inspection of text and medical media, tool use, and structured decisions. Results show that capability is compositional rather than uniform: models struggle most with interpreting volume-derived views and producing precise spatial coordinates, video results differ between ultrasound and surgical settings, and agentic success depends on both the clinical answer and workflow execution. MedStat-Bench therefore characterizes capability across medical data types, output formats, and tool-mediated workflows rather than relying on a single task or aggregate score.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.