Safety-Aware Value Alignment for Offline Safe Reinforcement Learning
Abstract
Offline safe reinforcement learning (OSRL) often extracts policies from learned value funcitons, but learned reward values can rank dataset-supported actions differently from the constrained objective. We introduce Safety-Aware Value Alignment (SAVA), which makes learned long-horizon cost information part of the Bellman target for reward value learning. Using cost values to identify high-cost regions, reshapes reward value learning when transitions enter these regions, allowing the resulting penalized Bellman targets to propagate to preceding state–action pairs and to reshape earlier action preferences before policy extraction. We theoretically establish sufficient conditions under which the penalized objective favors safe policies over unsafe ones and derive a bound on the reward gap to the optimal safe policy under value-estimation and policy-extraction errors. Empirically, controlled ablations show that the region-aware penalized reward value is the primary component responsible for constraint satisfaction. On the DSRL benchmark, SAVA achieves the highest normalized return on 27 of 32 tasks while satisfying the common normalized-cost safety criterion.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.