acceptodds
Under review as a conference paper at ICLR 2027

A Variance-Optimal Replay Framework for Safe Off-Policy Reinforcement Learning

Abstract

Existing off-policy safe reinforcement learning methods inherit replay sampling strategies developed for unconstrained settings. Consequently, replay sampling remains optimized for reward rather than safety, compromising constraint satisfaction. Methods that prioritize transitions using safety-related heuristics reshape the sampling distribution without correction, inadvertently shifting the very solution that the agent is meant to reach. In this work, we treat the replay sampling distribution as a control variable to improve the agent's safety and learning performance. We prove that importance-weighted replay sampling preserves the expected constrained policy update, while affecting only its variance. We further show that reducing this variance lowers the probability of unsafe updates during training. Building on this result, we derive a closed-form replay sampling strategy that achieves the minimum possible variance across minibatches and reduces safety violations. At convergence, it asymptotically approaches the optimal policy of the underlying Constrained Markov Decision Process. Our sampling method is compatible with a wide range of gradient-based safe reinforcement learning algorithms and demonstrates consistent improvements over existing replay sampling methods on the Safety Gymnasium benchmark suite.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.