The Road Not Ordered: Counterfactual Step Pricing for Reinforcement Learning in Sequential Medical Diagnosis
Abstract
Medical diagnosis is a sequential and cost-aware process: physicians order examinations, update their diagnostic reasoning as evidence accumulates, and make a final diagnosis once the evidence is sufficient. Reinforcement learning provides a natural framework for training sequential diagnostic agents. However, reinforcement learning in this setting faces two key challenges: reward blindness, where diagnostic paths with different levels of efficiency can receive the same score if they reach the same conclusion; and credit dilution, where one final score is shared across many steps, making the pivotal examination difficult to distinguish from routine or unnecessary ones. These challenges arise from assigning a single final-outcome reward to the entire trajectory. To address these issues, we propose a counterfactual reinforcement learning approach called Counterfactual Step Pricing (CoSteP). CoSteP focuses on decision-critical states where the policy is uncertain. At each selected state, it compares the chosen action with plausible counterfactual alternatives using short rollouts and a utility that considers both diagnostic accuracy and examination cost. A persistent rollout cache reuses compatible trajectories accumulated across training batches, reducing computation without expert step annotations or a learned critic. Experiments on MIMIC-IV, ClinicalBench, and a private hospital dataset show that CoSteP improves diagnostic accuracy while reducing both the number and cost of examinations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.