acceptodds
Under review as a conference paper at ICLR 2027

Latent Decision Flow Rewards for Clinical Guideline Reasoning

Abstract

Clinical practice guidelines encode evidence-based recommendations as conditional decision flows, which makes them a natural source of process supervision for training language models. Using this supervision is difficult for two reasons. A decision flow consistent with a guideline does not by itself determine the correct answer for a given patient, which we refer to as the path–answer gap, and several valid decision flows can support the same clinical decision, which we refer to as path multiplicity. Under these conditions, answer rewards provide no signal about the decision process, supervised training on annotated flows constrains the model to a single trajectory and reduces answer accuracy, and process rewards computed against one reference flow penalize alternative valid reasoning. We introduce GUIDER, a joint controller–policy optimization approach that learns guideline-adherent clinical reasoning while accounting for these challenges. The controller learns to generate plausible decision flows from guideline-derived supervision and provides process-level supervision to the policy, while the policy learns to generate clinical reasoning and final decisions. By aggregating process rewards over multiple controller-generated decision flows, \name accounts for path multiplicity, and by gating process rewards on answer correctness, it couples guideline adherence with correct final answers. Joint optimization allows process supervision to evolve with the policy instead of relying on a fixed reference trajectory. The policy is never conditioned on an oracle decision path, so GUIDER requires no access to guidelines at inference time. On MedGUIDE, \name with Qwen3-14B achieves 82.5 weighted accuracy and 73.5 guideline adherence, improving over the strongest clinical reinforcement learning baseline by 3.3 and 3.1 points, respectively. \name also improves over CRPO, MRPO, and FaithMed on the out-of-domain benchmarks AMEGA-LLM, RealMedQA, and CancerGUIDE with Qwen3-14B and Gemma4-12B backbones.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.