acceptodds
Under review as a conference paper at ICLR 2027

Beyond the Constraint: Probing Constrained RL Bottlenecks with Phasic Policy Gradient

Abstract

Constrained Reinforcement Learning (CRL) extends policy optimization to Constrained Markov Decision Processes (CMDPs), where agents must maximize reward while satisfying explicit cost budgets. Although existing on-policy CRL methods differ in how they enforce constraints, they typically follow the same training paradigm where collected experience is used once and then discarded. We investigate whether this practice underutilizes information contained in sampled CRL trajectories, which provide supervision through both reward and cost signals. To this end, we introduce a phasic training framework for CRL that separates policy optimization from a dedicated auxiliary phase. During this phase, retained rollouts are reused to refine reward and cost representations that are then distilled into the policy network. The framework acts as a modular wrapper and can extend existing CRL algorithms without modifying their constraint-handling mechanisms. Experiments on Safety Gymnasium show that our extension to CPO, CPPO-PID and TRPO-PID yields consistent improvements in sample efficiency and cumulative reward in environments with rich observation structure. This suggests that representation learning and data utilization are important and underexplored factors in on-policy CRL performance.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.