SHIFT: Benchmarking Clinical Computer-Use Agents over Full Patient Trajectories
Abstract
Computer-use agents operate software through its graphical interface and could work inside the electronic health record (EHR) systems clinicians use. Yet most medical AI benchmarks still rely on pre-curated text, images, or structured data, and evaluate isolated tasks rather than a continuous course of care. It therefore remains unclear whether agents can navigate clinical interfaces, maintain a patient's evolving state, and carry care from initial assessment through follow-up. Here we introduce SHIFT (Stateful Hospital Interaction for Full Trajectories), a benchmark for clinical computer-use agents over complete patient trajectories. SHIFT builds on a real EHR (OpenMRS) and image viewer (OHIF), connecting interviews, charting, orders, results, image interpretation, diagnosis, treatment, and follow-up to a shared patient state. We score agents along seven clinical capabilities, using case-specific dependency graphs to check the evidence supporting each scoring item while accepting clinically valid alternatives. Across nine models, GPT-6-Astra achieves the highest mean score of 53.82 out of 100. Five models score below 30. Common errors include incorrect patient information, misinterpreted test results, and neglected earlier findings. Scores are lower in completed six-patient sessions. By exposing where agents lose track of patient state and evidence, SHIFT quantifies the gap that must be closed before computer-use agents can safely operate real clinical software.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.