acceptodds
Under review as a conference paper at ICLR 2027

SAGA-CDP: Large Language Models Can Learn Guideline-Adherent Clinical Decision-Making from Synthetic Data

Abstract

Reliable clinical decision-making with large language models requires models to adhere to specified clinical guidelines throughout the decision-making process. In this work, we systematically investigate whether guideline adherence can be learned as a transferable capability from both real-world and synthetic guideline data. These two data sources present distinct trade-offs in guideline-adherence training. To enable structured synthetic data generation, we introduce SAGA-CDP, a data generation pipeline that produces fictional clinical guidelines, clinical decision pathways (CDPs), and case vignettes from shared underlying decision logic, yielding 2,000 fictional guidelines and 76,584 corresponding clinical cases. Extensive experiments demonstrate that real-world and fictional guideline data are both effective training sources, and that combining them achieves the strongest overall performance. Scaling analyses reveal distinct behaviors across the two data sources, with real-world guideline data continuing to benefit from more training cases, while further gains from fictional data depend more strongly on broader guideline coverage. Cross-regional real-world studies and controlled synthetic interventions provide evidence of context-memory conflict in guideline-adherent clinical decision-making. These findings establish real-world guideline data as a clinically grounded and viable resource for guideline-adherence training, while highlighting synthetic data as a scalable complement that increases the amount of training data and reduces reliance on real-world guideline knowledge.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.