A Variance-Optimal Replay Framework for Safe Off-Policy Reinforcement Learning
Abstract
Existing off-policy safe reinforcement learning methods inherit replay sampling strategies developed for unconstrained settings. Consequently, replay sampling remains optimized for reward rather than safety, compromising constraint satisfaction. Methods that prioritize transitions using safety-related heuristics reshape the sampling distribution without correction, inadvertently shifting the very solution that the agent is meant to reach. In this work, we treat the replay sampling distribution as a control variable to improve the agent's safety and learning performance. We prove that importance-weighted replay sampling preserves the expected constrained policy update, while affecting only its variance. We further show that reducing this variance lowers the probability of unsafe updates during training. Building on this result, we derive a closed-form replay sampling strategy that achieves the minimum possible variance across minibatches and reduces safety violations. At convergence, it asymptotically approaches the optimal policy of the underlying Constrained Markov Decision Process. Our sampling method is compatible with a wide range of gradient-based safe reinforcement learning algorithms and demonstrates consistent improvements over existing replay sampling methods on the Safety Gymnasium benchmark suite.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.