acceptodds
Under review as a conference paper at ICLR 2027

MultiKnee: A Multi-Sequence Multi-View MRI Benchmark for Evidence-Chain Knee Diagnosis

Abstract

Medical imaging benchmarks for multimodal large language models (MLLMs) predominantly rely on single-image inputs or coarse-grained labels, as fine-grained annotations aligned across sequences and views are prohibitively expensive. We introduce **MultiKnee**, a hierarchically structured multi-sequence, multi-view knee MRI benchmark with fine-grained annotations at scale: 11,575 studies, 104,175 sequence-view units, and 821,823 VQAs across six anatomical systems, 71 clinical attributes, and 34 diagnostic categories. VQAs are organized as **explicitly linked Observation-Evidence-Diagnosis chains** that trace a clinical reasoning pathway: from perceiving images across sequences and views, to interpreting them within anatomical context, and reaching the final diagnosis. Across 20 model configurations, controlled-observation analysis reveals substantial differences in how models utilize complementary sequence and view information. Stage-wise analysis further reveals that models' reasoning bottlenecks lie primarily in evidence interpretation and diagnosis rather than observation. MultiKnee thus provides a large-scale infrastructure for multi-sequence, multi-view MRI research and a framework for understanding how MLLMs reason from observations to diagnoses.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.