CareLoop: Measuring the Gap Between Knowing Medicine and Practicing It in the Context of Patients' Lives
Abstract
Medical benchmarks usually ask whether an AI system knows the right answer. Patient-facing care also requires discovering incomplete or incorrect information, understanding a person's beliefs and constraints, adapting feasible plans, verifying execution, incorporating delayed evidence, and retaining responsibility for unresolved risk. We introduce CareLoop, a clinical-world sandbox that measures this gap between knowing medicine and practicing it in the context of patients' lives. Its 120 deidentified EHR-derived case contracts distribute state across clinical truth, patient and family beliefs, records, and executed actions; prospectively authored frictions create controlled behavioral forks with observable downstream consequences. Ten models generated 1,200 trajectories. Four blinded LLM Judges produced 4,800 complete-context assessments, and 139 physicians produced 6,000 structured assessments. GPT-5.6 Sol ranked first on the Friction Capability Composite (FCC, 0.824) and its safety-gated counterpart C-RWR (0.753). Across 15,537 realized high-order opportunity assessments, models acted in 0.935, changed the trajectory meaningfully in 0.745, and completed the opportunity effectively in 0.697; hidden-state discovery was the largest shared bottleneck. Individual assessments varied, yet disjoint 60-case halves recovered stable model rankings (median Spearman for FCC and for C-RWR), and aggregate physician and four-Judge rankings converged ( and ). Matched model–Judge analyses showed no uniform self-favoring pattern. Together, these results show how execution-grounded evaluation can extend from directly testable state change to consequential action under partial observability and limited control: the benchmark makes the patient's lived context causally operative and aggregates structured observations over broad case coverage. Anonymous code and data: https://anonymous.4open.science/r/careloop-iclr2027-6e55. The simulation does not establish clinical effectiveness or deployment safety.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.