MedSensus: Benchmarking Medical Reasoning in Omni-Modal Models
Abstract
Clinical reasoning rarely relies on a single source of evidence. Clinicians combine complementary observations to understand the same problem from different perspectives. Although modern models can now process heterogeneous inputs together, publicly accessible medical datasets with reliable clinical annotation remain scarce. Most existing benchmarks still focus on individual modalities or narrower pairings. We introduce MedSensus, a medical benchmark built from publicly available resources and reviewed by clinical experts. MedSensus comprises 9,539 evaluation instances organized into 1,865 evidence groups that share visual and acoustic evidence. Questions may have one or multiple correct answers. This structure supports question accuracy and group accuracy, with the latter requiring all related questions to be answered correctly. The benchmark also provides a factorial evaluation suite that selectively withholds the image, audio, or question stem, enabling estimation of the conditional ablation effect of each input source. Across fifteen representative models, the best model under direct prompting reaches 63.84% question accuracy and 17.04% group accuracy, compared with the clinician reference of 91.12% and 72.97%. These results show that accuracy on individual questions does not imply consistent reasoning across clinical evidence. By combining expert validation with controlled diagnostic evaluation, MedSensus provides a basis for studying how medical models integrate heterogeneous evidence.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.