acceptodds
Under review as a conference paper at ICLR 2027

SafeOT: A Unified Safety-Price View of Safe Reinforcement Learning

Abstract

Safe reinforcement learning (safe RL) is commonly formulated in constrained Markov decision processes (CMDPs) with explicit constraints on expected safety costs, or optimized through scalarized penalty and Lagrangian objectives. Although both address the same reward–safety trade-off, they are usually treated as distinct families of methods rather than within a single operational framework, and common multi-constraint extensions either tune one Lagrange multiplier per cost or rely on scalarized penalties, which require per-constraint weighting and are prone to oscillation as the number of costs grows. This makes it difficult to jointly represent and control multiple safety constraints with distinct cost functions and bounds. We propose SafeOT, which recasts safe policy optimization as a budget- constrained optimal-flow problem over the transition graph of a CMDP. This representation places penalty-based and constraint-based safe RL on one spectrum and handles all constraints in a single projection. Coupling this projection with a PPO-like per-iteration policy update yields SafeOT, a GPU-efficient algorithm whose projection takes under 0.2 seconds per update. On 16 Safety-Gymnasium and Bullet-Safety-Gym tasks, SafeOT reduces cumulative constraint violation relative to TRPO-Lagrangian by 63% and 70%, respectively, lowering it on every task with comparable return on Safety-Gymnasium. With six to eight per-joint budgets on MuJoCo robots, its unclipped variant also incurs less violation than Lagrangian and PID updates.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.