Beyond the Constraint: Probing Constrained RL Bottlenecks with Phasic Policy Gradient
Abstract
Constrained Reinforcement Learning (CRL) extends policy optimization to Constrained Markov Decision Processes (CMDPs), where agents must maximize reward while satisfying explicit cost budgets. Although existing on-policy CRL methods differ in how they enforce constraints, they typically follow the same training paradigm where collected experience is used once and then discarded. We investigate whether this practice underutilizes information contained in sampled CRL trajectories, which provide supervision through both reward and cost signals. To this end, we introduce a phasic training framework for CRL that separates policy optimization from a dedicated auxiliary phase. During this phase, retained rollouts are reused to refine reward and cost representations that are then distilled into the policy network. The framework acts as a modular wrapper and can extend existing CRL algorithms without modifying their constraint-handling mechanisms. Experiments on Safety Gymnasium show that our extension to CPO, CPPO-PID and TRPO-PID yields consistent improvements in sample efficiency and cumulative reward in environments with rich observation structure. This suggests that representation learning and data utilization are important and underexplored factors in on-policy CRL performance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.