acceptodds
Under review as a conference paper at ICLR 2027

A Unified Bellman Operator for Safety-Critical Reinforcement Learning

Abstract

Reinforcement learning in safety-critical domains requires maximizing task performance while strictly adhering to safety constraints. Existing safe reinforcement learning paradigms *typically* force a trade-off: they either require *a priori* knowledge to provide strict safety guarantees (e.g., safety filters), or they enable joint learning but only satisfy safety constraints on average. In this work, we propose a novel Bellman operator that unifies performance and safety objectives into a *joint* value function. We show that temporal difference learning with the joint Bellman operator converges under a two-timescale stochastic approximation framework. On the fast timescale, the *safety value* of the learning joint policy is estimated, while the *joint value* is estimated on the slow timescale. Convergence is ensured by formulating the limiting dynamics as an *occupation-averaged* differential inclusion, and showing that it asymptotically converges to a set of limiting optimal safety-constrained task value functions. *Theoretically*, once converged, the resulting optimal policy maximizes task return while maintaining safety at all times. Empirical evaluations on continuous control tasks with *neural approximations* demonstrate stable convergence with near-zero safety violations at test time.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.