acceptodds
Under review as a conference paper at ICLR 2027

Do We Need Cost Critics? Local Cost Relabeling for Offline Safe Reinforcement Learning

Abstract

Offline safe reinforcement learning commonly relies on learned cost-to-go functions, feasibility models, or adaptive penalty mechanisms to enforce cumulative-cost constraints. We ask how much of this machinery is necessary when safety violations are directly observed in the offline data. We study a simple alternative, RELABEL(COST), that retains the full dataset, replaces the reward on positive-cost transitions by a fixed negative value derived from the reward scale, and then trains a behavior-regularized offline RL policy. The method uses no cost critic, no feasibility estimator, and no safety-specific penalty tuning. In backbone-matched continuous-control experiments, RELABEL(COST), consistently satisfies the safety budget and achieves near-zero cost on most tasks while remaining competitive with, and often outperforming, adaptive additive penalties and cost-critic-based relabeling. These comparisons also reveal an important role for strong behavior regularization, with local cost supervision identifying undesirable transitions while regularization limits policy improvement to regions supported by the offline data. Our analysis characterizes when fixed local penalties align with zero-cost policy optimization, explains how delayed cost events propagate through ordinary reward value learning, and shows that safe behavior can in principle be recovered even when complete trajectories in the dataset are infeasible. Together, the results suggest that direct local cost supervision combined with strong offline regularization can provide a simple and robust alternative to explicit cumulative-cost modeling in offline safe RL.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.