TRACE: Transparent Reasoning through Aligned Concept Explanations in Clinical Reinforcement Learning
Abstract
When an offline reinforcement learning (RL) policy recommends a treatment that differs from observed clinical practice, the disagreement alone cannot reveal whether the departure reflects clinically useful personalization or limitations in the data or learning objective. We introduce Transparent Reasoning through Aligned Concept Explanations (TRACE), a framework that explains and directly compares a learned RL policy with an estimated behavior policy fitted to recorded treatment actions. TRACE uses a shared encoder and sparse autoencoder to define a common concept space, then conditions a shared LLM-based explainer on the patient state and policy-relevant concepts while withholding the selected action. Teacher supervision and proximal policy optimization promote hidden-action recovery, concept consistency, state grounding, and resistance to action leakage. We evaluate TRACE on three MIMIC-III/IV tasks, including renal replacement therapy, noninvasive ventilation, and vasopressin initiation, where the policies disagree on 11.9%–33.4% of held-out states. TRACE achieves the highest hidden-action recovery accuracy and macro-F1 among the evaluated explanation methods for both policies. On disagreement states, each policy's action is recovered more accurately from its own explanation than from the other policy's explanation. Concept perturbations reduce learned-policy recovery across all tasks, while policy-side interventions show that cited concepts are more influential than uncited or random concepts, most clearly for RRT. These results show that TRACE enables faithful, concept-grounded comparison of learned and behavior policies, providing a basis for pre-deployment auditing and targeted clinical review of disagreements.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.