When Zero Cost Does Not Mean Safe: Response Delay and Policy Delivery in Resource-Constrained Reinforcement Learning
Abstract
Safe reinforcement learning under a finite training budget can fail even when standard constraint mechanisms are in place: resource degradation may raise future risk before positive cost is observed, constraint response may be delayed, and feasible policies found during training may be lost before delivery. We analyze these failure stages, separating non-detection from post-detection non-recovery via multiplier first-passage analysis and decomposing final policy delivery into candidate opportunity and selection loss. Guided by this analysis, we propose Enhanced Resource-Coupled Trust Region Policy Optimization with Shadow Generation (ERC-TRPO-SG), which integrates resource-aware candidate generation, conditional shadow repair, and policy retention. Controlled interventions and factorial ablations show that resource state encodes future-risk information beyond instantaneous constraint cost and that policy retention accounts for most of the improvement in feasible-policy delivery. Across three resource-coupled domains, ERC-TRPO-SG yields mean-feasible policies for all 20 training seeds in each domain while maintaining competitive task returns. We further conduct matched 10M-interaction evaluations on nine public Safety-Gymnasium tasks. While CSPO recovers more quickly from early constraint violations, our ERC-U base variant exhibits far fewer late-stage violations, lower mean constraint cost on 7/9 tasks, and mean-feasible frozen policies on 8/9 tasks versus 6/9 for CSPO. These results distinguish initial recovery, persistent safety, and final policy delivery as separate finite-budget objectives.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.