acceptodds
Under review as a conference paper at ICLR 2027

MedSensus: Benchmarking Medical Reasoning in Omni-Modal Models

Abstract

Clinical reasoning rarely relies on a single source of evidence. Clinicians combine complementary observations to understand the same problem from different perspectives. Although modern models can now process heterogeneous inputs together, publicly accessible medical datasets with reliable clinical annotation remain scarce. Most existing benchmarks still focus on individual modalities or narrower pairings. We introduce MedSensus, a medical benchmark built from publicly available resources and reviewed by clinical experts. MedSensus comprises 9,539 evaluation instances organized into 1,865 evidence groups that share visual and acoustic evidence. Questions may have one or multiple correct answers. This structure supports question accuracy and group accuracy, with the latter requiring all related questions to be answered correctly. The benchmark also provides a factorial evaluation suite that selectively withholds the image, audio, or question stem, enabling estimation of the conditional ablation effect of each input source. Across fifteen representative models, the best model under direct prompting reaches 63.84% question accuracy and 17.04% group accuracy, compared with the clinician reference of 91.12% and 72.97%. These results show that accuracy on individual questions does not imply consistent reasoning across clinical evidence. By combining expert validation with controlled diagnostic evaluation, MedSensus provides a basis for studying how medical models integrate heterogeneous evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.