acceptodds
Under review as a conference paper at ICLR 2027

Survival Policy Learning as Probabilistic Inference

Abstract

Learning optimal treatment policies that maximize survival probabilities from longitudinal observational data is a central problem in causal inference for medicine. Existing applications of reinforcement learning (RL) to longitudinal medical data typically adapt standard RL algorithms under a Markov decision process formulation and rely on heuristic rewards specified by clinicians or auxiliary models, thereby weakening the scientific validity and interpretability of the learned policies. In this study, we formulate causal survival policy learning within the framework of control as inference. This formulation yields a principled dense reward which connects reward design directly to survival quantities and induces a behavior-anchored, KL-regularized policy learning objective for offline observational data. We further use an autoregressive model to learn state representations that approximate the Markov property of the sequential decision process. We evaluate the proposed method using synthetic data and semi-synthetic data based on real-world chronic disease epidemiological study, tracking individual health trajectories over a 30-year horizon. Across experiments, the proposed survival-derived reward provides an effective alternative to sparse survival rewards for survival policy learning.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.