acceptodds
Under review as a conference paper at ICLR 2027

When Zero Cost Does Not Mean Safe: Response Delay and Policy Delivery in Resource-Constrained Reinforcement Learning

Abstract

Safe reinforcement learning under a finite training budget can fail even when standard constraint mechanisms are in place: resource degradation may raise future risk before positive cost is observed, constraint response may be delayed, and feasible policies found during training may be lost before delivery. We analyze these failure stages, separating non-detection from post-detection non-recovery via multiplier first-passage analysis and decomposing final policy delivery into candidate opportunity and selection loss. Guided by this analysis, we propose Enhanced Resource-Coupled Trust Region Policy Optimization with Shadow Generation (ERC-TRPO-SG), which integrates resource-aware candidate generation, conditional shadow repair, and policy retention. Controlled interventions and factorial ablations show that resource state encodes future-risk information beyond instantaneous constraint cost and that policy retention accounts for most of the improvement in feasible-policy delivery. Across three resource-coupled domains, ERC-TRPO-SG yields mean-feasible policies for all 20 training seeds in each domain while maintaining competitive task returns. We further conduct matched 10M-interaction evaluations on nine public Safety-Gymnasium tasks. While CSPO recovers more quickly from early constraint violations, our ERC-U base variant exhibits far fewer late-stage violations, lower mean constraint cost on 7/9 tasks, and mean-feasible frozen policies on 8/9 tasks versus 6/9 for CSPO. These results distinguish initial recovery, persistent safety, and final policy delivery as separate finite-budget objectives.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.