Beyond Lagrangian Methods for Robust Peak-Cost Constrained Reinforcement Learning
Abstract
We study robust peak-cost constrained reinforcement learning, where the goal is to maximize reward while bounding the largest cost incurred along a trajectory – unlike the constrained MDP (CMDP), which only keeps the expected cumulative cost below a threshold. We first show that the problem is exactly equivalent to a CMDP on a state augmented with the running maximum cost, which restores zero duality gap but requires history-dependent policies, yields a cost signal that is active only at a new peak, and makes worst-case evaluation conservative. We therefore keep the reachability-constrained (RCRL) formulation on the original state space, for which strong duality can fail, and optimize a surrogate rather than a Lagrangian, estimating worst-case values through an integral probability metric. Our main contribution is the first convergence analysis for this surrogate. We prove that every iterate stays -safe from a safe start which the primal-dual-based approach can not guarantee, and that a stationary point of the exact surrogate is reached. Under a nondegeneracy condition, we further obtain a globally -optimal and -safe policy in iterations from a safe start, and otherwise. Experiments on continuous-control benchmarks, including RCRL, augmented CMDP and robust CMDP baselines, show that the method sustains reward while enforcing peak-cost safety under dynamics perturbations.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.