CohortWorld: Aligning Real-World Data-Driven Diagnostics and Language Models
Abstract
Medical language models (LMs) have shown great potential in complex real-world healthcare tasks since they can act as agents in interactive virtual clinical environments with verifiable rewards as trustworthy feedback. However, the accuracy of data-driven diagnostics in clinic practice is challenged by label quality due to restrictive data use agreements, synthetic yet access-gated health environment, and overly narrow health context, which leads to the absence of real-world evidence (RWE) verifying LMs to avoid hallucination in medical LMs. To bridge this gap, we present CohortWorld, an in-silico environment generated from large-scale UK Biobank (UKB) cohort in which each state is a real-world cohort anchored in disjoint histories of disease, procedure, and medication. CohortWorld is an reinforcement learning (RL) suite based on disjoint anonymous cohorts (=29,884) with real-world data gathered from 500k UKB subjects for evaluating LMs on disease diagnosis and simulating RWE reasoning as a human doctor. Our empirical testing demonstrates that existing closed- and open-source LMs severely underperform on the RWE-calibrated benchmark, hovering at or below an AUROC of 0.5 with suboptimal slope metrics for disease risk ranking. In contrast, Qwen3.5-9B outperforms Sonnet 5 using our CohortWorld framework, demonstrating the benefits of RWE matching and counterfactual risk modification. Together, CohortWorld advances the trustworthiness of artificial diagnosis and enables LMs to reason like a human doctor based on verifiable RWE in a releasable and actionable in-silico environment.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.