VasoDPO: Aligning Treatment Actions with Patient States through State-Fixed Preference Construction
Abstract
Clinical treatment generation aims to tailor actions to patients, yet retrospective electronic health records provide limited preference supervision over alternative actions under the same patient state. We study this problem through retrospective single-step generation of normalized norepinephrine actions for patients with sepsis in intensive care units. We propose VasoDPO, a state-fixed Direct Preference Optimization (DPO) framework that anchors chosen responses to recorded actions and uses mean arterial pressure (MAP) to construct rejected actions under the same patient state. After minimum action-gap filtering, VasoDPO upweights retained pairs contrasting a nonzero recorded action with a zero rejected action, motivated by near-zero concentration observed after supervised fine-tuning. Because favorable action-replacement train-on-synthetic, test-on-real (TSTR) scores can coexist with near-zero action concentration when state features are retained, we jointly evaluate TSTR with direct diagnostics of near-zero concentration, directional violations, and MAP sensitivity. Across five training seeds on a fixed split, VasoDPO achieves higher mean TSTR AUROC and AUPRC and lower mean near-zero and directional-violation rates than a matched DPO baseline using the same preference pairs with uniform weights. Matched ablations support the contributions of MAP-guided construction and omission weighting, while state–pair shuffling degrades both TSTR and action diagnostics. Controlled MAP sweeps show that the mean generated action varies with MAP while other inputs remain fixed.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.