acceptodds
Under review as a conference paper at ICLR 2027

On the Reliability of Value Estimation for Safe Exploration

Abstract

Safe exploration requires reinforcement learning (RL) agents to effectively trade off task performance against the risk of constraint violations during learning. Constrained RL algorithms use reward and cost value estimates to guide this trade-off, while their theoretical guarantees are stated in terms of the true value functions that these estimates approximate. Yet the reliability of the cost critic is rarely examined directly. We show that cost critics exhibit substantially poorer agreement with Monte Carlo reference values than reward critics. The regression problem for the cost critic spans a spectrum, from heavily bootstrapped targets that are stable but inherit the critic's own errors, to long-horizon targets that contain more information about future cost but are high variance and difficult to generalize from. Neither end yields a critic that correlates well with the reference cost value, though for different reasons. Because cost-critic errors do not exhibit a systematic bias, existing bias-control methods do not recover value fidelity and instead primarily make policy updates more conservative; regularization likewise fails to improve the overall safety–performance trade-off. We therefore explore successor representations as an alternative approach to learning cost value functions and show that they provide a promising direction for safer and more sample-efficient exploration.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.