ClinTRACE: Benchmarking Clinical Reasoning as Evidence Accumulates
Abstract
AI-assisted clinical decision support requires large language models (LLMs) to reason and make decisions as evidence accumulates, not merely predict final diagnoses. However, existing benchmarks rarely provide expert reasoning traces aligned with the evidence available at each stage. To address this gap, we introduce ClinTRACE, a longitudinal clinical reasoning benchmark built from live, multi-expert discussions of real-world cases. Unlike retrospective rationales, these discussions capture physicians' reasoning before subsequent findings are revealed. Using a physician co-designed taxonomy and an audited data-extraction pipeline, we structure these discussions into time-aligned trajectories that link available evidence to expert interpretations, hypothesis updates, and proposed actions. The dataset includes 1,464 public cases with 84,044 atomic reasoning steps, plus 184 non-public cases. We formulate five complementary tasks to evaluate longitudinal clinical reasoning: progressive diagnosis and reasoning, action planning, hypothesis tracking, evidence-hypothesis relation prediction, and clinical statement typing. Evaluation of 11 general-purpose and medical LLMs shows that strong diagnostic performance coexists with incomplete coverage of documented expert reasoning. The leading model achieves 76.7% top-5 diagnostic accuracy but only 33.3% reasoning comprehensiveness. Supervised fine-tuning improves longitudinal reasoning capabilities and transfers beyond ClinTRACE, with diagnostic gains on two external benchmarks. Together, ClinTRACE supports evaluation and learning of clinical reasoning as evidence accumulates, beyond final-answer accuracy. We will release the public dataset and retain the non-public cases as a protected evaluation set.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.