SafeOT: A Unified Safety-Price View of Safe Reinforcement Learning
Abstract
Safe reinforcement learning (safe RL) is commonly formulated in constrained Markov decision processes (CMDPs) with explicit constraints on expected safety costs, or optimized through scalarized penalty and Lagrangian objectives. Although both address the same reward–safety trade-off, they are usually treated as distinct families of methods rather than within a single operational framework, and common multi-constraint extensions either tune one Lagrange multiplier per cost or rely on scalarized penalties, which require per-constraint weighting and are prone to oscillation as the number of costs grows. This makes it difficult to jointly represent and control multiple safety constraints with distinct cost functions and bounds. We propose SafeOT, which recasts safe policy optimization as a budget- constrained optimal-flow problem over the transition graph of a CMDP. This representation places penalty-based and constraint-based safe RL on one spectrum and handles all constraints in a single projection. Coupling this projection with a PPO-like per-iteration policy update yields SafeOT, a GPU-efficient algorithm whose projection takes under 0.2 seconds per update. On 16 Safety-Gymnasium and Bullet-Safety-Gym tasks, SafeOT reduces cumulative constraint violation relative to TRPO-Lagrangian by 63% and 70%, respectively, lowering it on every task with comparable return on Safety-Gymnasium. With six to eight per-joint budgets on MuJoCo robots, its unclipped variant also incurs less violation than Lagrangian and PID updates.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.