acceptodds
Under review as a conference paper at ICLR 2027

Toward Virtual Patients: A Benchmark for Forecasting Future Radiology Reports

Abstract

Retrospective electronic health records (EHRs) contain outcomes only for the tests clinicians historically ordered, so interactive clinical agents cannot be trained or evaluated on queries that depart from recorded trajectories. We study an evaluable surrogate—query-conditioned clinical world modeling: given a patient's leakage-controlled EHR context and a requested imaging study, generate the study's report, scored by its structured findings under a frozen extractor. Building on MIMIC-IV and MIMIC-IV-Note, we align longitudinal clinical events with 117K radiology reports across four modalities and pose 1,500 sealed masked-observed queries (every target study was actually performed, so every query is gradable). Evaluating frontier LLMs, structured supervised pipelines, and a budgeted LLM-policy search over model design, we find a robust dissociation: end-to-end LLMs reach micro-F1 while structured predictors over a prior-imaging finding-state representation plateau at , invariant to search budget, policy strength, and rendering strategy. We trace the plateau to the supervision itself: report labels measure radiologists' mention behavior, not patient state—a finding documented as present goes unmentioned in the next report of the time even within seven days, nearly independent of elapsed time. Under an omission-tolerant reading of the same predictions—forgiving a predicted finding that prior reports document and the target report merely leaves unmentioned—all systems gain 0.10–0.12 and the verdict changes: a structured pipeline that significantly trails the best end-to-end LLM under the mention convention ties it under that reading—the evaluation convention, not only the model, determines the conclusion. We formalize the task as state estimation composed with mention selection and intervene on both factors, obtaining a double dissociation: mention-trained models win under the mention convention while densely state-supervised models win under the omission-tolerant reading—where a dense-supervised world model nearly matches the best end-to-end LLM—and, at matched context budget, structured state supplements lift an end-to-end report generator by . Our model-design search runs under an integrity-by-construction protocol: the LLM policy acts in a typed, budgeted action space with the sealed test structurally unreachable; one committed configuration is scored once by an external evaluator. Validation-to-sealed gaps grow with search intensity while sealed scores stay flat, and policy behavior (early saturation, cost-insensitivity, avoidance of semantic actions) is reported as part of the evaluation. We release the benchmark, findings schema, extraction pipeline, environment, and evaluation harness.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.