acceptodds
Under review as a conference paper at ICLR 2027

Confidence-aware chance-constrained policy optimization

Abstract

In many real-world applications, reinforcement learning (RL) agents are expected to balance reward maximization and safety considerations. Constrained RL can address this issue by maximizing expected cumulative return while satisfying safety-related constraints. While RL algorithms with expected cumulative cost constraints have been extensively investigated, these expectation-based formulations provide a limited characterization of risk and do not directly capture the probability of safety violations. In this work, we propose a confidence-aware chance-constrained policy optimization (Chance-CPO) method that keeps the probability of constraint violation below a prescribed threshold at a user-specified confidence level. The proposed approach establishes a connection between expectation-based constraints and chance constraints, and uses expectation-related gradients to facilitate the satisfaction of chance constraints. By reducing an upper bound on the violation probability, Chance-CPO accounts for both the expected cumulative cost and the probability of unsafe events. We demonstrate the effectiveness of the proposed method on simulated robot tasks for which safety constraints are important.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.