acceptodds
Under review as a conference paper at ICLR 2027

Hybrid RL for Optimal Dynamic Treatment Regimes: Impact of Unmeasured Confounding

Abstract

Hybrid Reinforcement Learning (RL), where an agent learns from both an offline dataset and online explorations in an unknown environment, has attracted significant recent interest. Existing hybrid RL studies have largely focused on two cases for the offline data-generating process, either imposing no restriction on the degree of unmeasured confounding or assuming that such confounding is entirely absent. Many applications, however, lie between these extremes, with unmeasured confounding present but bounded in magnitude. In this paper, we investigate how the degree of unmeasured confounding in the offline data impacts online learning for optimal dynamic treatment regimes. We show that the offline data never hurt the guarantee of the purely online setting. Moreover, the guarantee improves monotonically as the assumed degree of unmeasured confounding decreases. When the offline data are fully unconfounded, the guarantee is at least as strong as those in the purely online, unrestricted, and bounded unmeasured confounding settings.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.