Taming Phantom Costs: Incremental Lower Bounds for Constrained Reinforcement Learning
Abstract
Constrained Reinforcement Learning (CRL) promises to improve the safety and efficiency of Reinforcement Learning under real world assumptions by replacing the need to tune uniterpretable reward weights with interpretable constraints. However, CRL comes with its own exploration challenges: Cost estimators suffer from random initialization errors and bad generalization which produces spuriously high cost estimates. Since high costs states are, by design, underexplored, those errors never get corrected and cause hallucinated “phantom costs” that needlessly constrain the CRL policy. Existing EM-style surrogate methods address this via monotone safety improvement, but their convergence speed is entangled with the scale of the Q-function through a manually tuned hyperparameter, leading to slow and brittle feasibility recovery as the optimization over- and under- estimates the local uncertainty. We propose ILB-CRL (Incremental Lower Bound CRL), which replaces the hyperparameter with an automatically computed trust coefficient derived from a lower confidence bound (LCB) on the optimal safety improvement . This formulation bridges the gap between incremental improvements, hard projections, and statistical reliability of the cost estimates. We prove geometric convergence under very limited assumptions and demonstrate either above average or state of the art result on a large set of safety-gymnasium environments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.