Joint-risk Aware Constrained Risk-sensitive Reinforcement Learning
Abstract
Maximizing rewards under various cost constraints is a frequent requirement when applying reinforcement learning (RL) to domains such as robotics and energy systems. Since expected-cost constraints alone cannot prevent rare but severe violations, constrained risk-sensitive RL is introduced to address this limitation through constraints on the risk of cost returns. However, because these constraints assess each cost independently using marginal distributions, they enforce overly conservative trade-offs by penalizing physically impossible events. We propose JRCPO, which preserves expected-cost constraints while replacing disjoint marginal risk constraints into a joint risk constraint over the reward-cost distribution. Utilizing the coherent duality of the Conditional Value at Risk (CVaR), JRCPO evaluates the CVaR of a scalarized objective along the Lagrange multiplier-directed trade-off path. This approach captures the joint distribution and considers only physically realizable events, resulting in an efficient reward-cost trade-off. Consequently, we demonstrate via evaluations on multi-cost tasks that the proposed method achieves higher reward returns than methods with cost-wise independent marginal risk constraints, and lower cost returns than expectation-based methods.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.