acceptodds
Under review as a conference paper at ICLR 2027

Zero-Sum Bias Corrections for Confounding by Indication in Offline Reinforcement Learning

Abstract

Although offline reinforcement learning promises more personalized treatment regimes in clinical care, challenges remain when action- and outcome-relevant information goes unrecorded in the data used to train policies. Standard value-based offline reinforcement learning methods converge to the observational value of actions and conflate the reason the action was chosen with its effect. Clinical data commonly suffers from confounding by indication, where patients with unrecorded indicators of disease severity both receive more aggressive treatment and experience worse outcomes. This biases naive learners, undervaluing life-saving treatments while overvaluing inaction. Many confounding-robust offline reinforcement learning methods impose bounds that further penalize treatment, while others require strict identification conditions that may be false in practice. In both cases, the learned policy can collapse and fail to provide any patient-specific targeting at all. We introduce an offline reinforcement learning algorithm to correct the training targets under static confounders. Our correction is derived from an omitted-variable-bias factorization and approximates the bias by the product of two components estimated from the full trajectory history, with the direction of the correction given by the data as the sign of a within-trajectory covariance. As per-action biases at a state provably cancel, we constrain our correction to be zero-sum across actions so that value is redistributed rather than uniformly suppressed. Across three synthetic medical environments, one fit to real clinical dynamics, our method produces higher-value policies than competing deconfounding and offline reinforcement learning approaches. Where conservative baselines collapse, our method continues to differentiate between patients.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.