acceptodds
Under review as a conference paper at ICLR 2027

MedXS: Are Medical Multimodal Agents Reasoning Across Sources, or Merely Landing on the Right Answer?

Abstract

Accurate clinical decisions rarely rest on a single image. They interleave visual inspection, retrieval over case specific structured records, and corroboration against external medical knowledge. First generation medical agent benchmarks judge outcome alone, and a second adds process aware scoring of tool calls, yet both share one blind spot. They confine tools to a single evidence modality, never treating the structured record as co equal with imaging, and their process metrics check whether calls are well formed rather than whether cross source evidence is used, so a trajectory grounded on one source still scores fully. We introduce MedXS, a process verified benchmark that instantiates imaging, structured record, and medical knowledge as three co equal tool families over 2,000 clinician annotated tasks, with eight question types stratified into three levels by the number of sources fused and grounded on 6,000+ evidence checkpoints, scored by an unweighted, tier increasing suite of six metrics spanning outcome, process, and cross source integration. Across frontier agents, the strongest reaches only 68.2% average result yet a 97.4% workflow score, while its Level 3 coverage collapses to 28.8% and knowledge tools drop to 2.8% of all calls. This exposes tool tunnel vision, a persistent gap between correct answers, valid tool ordering, and real cross source integration that stays invisible to both outcome and call level scoring.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.