Do We Need Cost Critics? Local Cost Relabeling for Offline Safe Reinforcement Learning
Abstract
Offline safe reinforcement learning commonly relies on learned cost-to-go functions, feasibility models, or adaptive penalty mechanisms to enforce cumulative-cost constraints. We ask how much of this machinery is necessary when safety violations are directly observed in the offline data. We study a simple alternative, RELABEL(COST), that retains the full dataset, replaces the reward on positive-cost transitions by a fixed negative value derived from the reward scale, and then trains a behavior-regularized offline RL policy. The method uses no cost critic, no feasibility estimator, and no safety-specific penalty tuning. In backbone-matched continuous-control experiments, RELABEL(COST), consistently satisfies the safety budget and achieves near-zero cost on most tasks while remaining competitive with, and often outperforming, adaptive additive penalties and cost-critic-based relabeling. These comparisons also reveal an important role for strong behavior regularization, with local cost supervision identifying undesirable transitions while regularization limits policy improvement to regions supported by the offline data. Our analysis characterizes when fixed local penalties align with zero-cost policy optimization, explains how delayed cost events propagate through ordinary reward value learning, and shows that safe behavior can in principle be recovered even when complete trajectories in the dataset are infeasible. Together, the results suggest that direct local cost supervision combined with strong offline regularization can provide a simple and robust alternative to explicit cumulative-cost modeling in offline safe RL.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.