PhysicianUIBench: Benchmarking Computer-Use Agents through a High-Fidelity EHR Interface
Abstract
Computer-use agents offer a way to automate clinical workflows through the electronic health record (EHR) interfaces physicians use in practice. Yet existing evaluations provide limited coverage of end-to-end workflows that require both clinical reasoning and multi-step EHR interaction. We introduce PhysicianUIBench, a benchmark for physician-annotated, real-world clinical workflows in a high-fidelity, interactive EHR simulation with an Epic-inspired interface. The benchmark comprises 100 long-horizon tasks and 772 checkpoints. Each task requires an agent, observing only screenshots and acting through mouse and keyboard, to locate evidence across a longitudinal patient chart, reach a clinical decision, and carry it through the interface by placing orders and writing the note, so that chart grounding, clinical reasoning, and EHR execution are exercised within a single workflow. Evaluation combines deterministic checks of application state and interaction logs with rubric-based assessment of clinical documentation. Among all evaluated agents, the highest task success rate is only 35%. These results reveal substantial limitations in current computer-use agents' ability to complete end-to-end clinical workflows reliably. PhysicianUIBench provides a testbed for evaluating how agents integrate clinical reasoning with EHR interface execution.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.