OmniClinic: Evaluating Active Evidence Acquisition in a Multimodal Clinical Sandbox
Abstract
Medical artificial intelligence holds great promise for improving clinical care, making rigorous and realistic evaluation essential for its safe deployment. However, existing benchmarks assess clinical capability mainly through text-only dialogue, overlooking the proactive multimodal evidence acquisition required in real-world diagnosis. To address this gap, we introduce OmniClinic, an interactive multimodal benchmark in which autonomous doctor agents acquire diagnostic evidence through clinical dialogue, diagnostic tests, and tool-mediated physical examinations. Agents operate within a finite clinical action space, selecting appropriate actions and specifying the parameters required for execution. We further introduce OmniDoctorX, which coordinates evidence-acquisition planning, spatial execution, and diagnostic review to support proactive clinical action. Across the evaluated omni-models, multimodal interaction improve diagnostic accuracy over text-only interaction. However, models struggle with evidence acquisition and spatial coordination, revealing a gap between multimodal understanding and autonomous interaction. OmniDoctorX improves mean proactive diagnostic accuracy from (86.36 vs. 77.84), demonstrating the value of coordinated acquisition and reasoning. Expert evaluation further supports the data quality of the benchmark, strengthening the reliability as a sandbox for autonomous multimodal diagnosis.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.