DRESS: Auditing Reward-Preserving Safety-Cost Amplification in Offline RL
Abstract
Offline safe reinforcement learning is commonly evaluated under nominal dynamics and monitored through task return or binary constraint violations. Such evaluation can miss a critical deployment failure mode: a small environment shift may preserve task performance while substantially amplifying safety cost. We introduce DRESS (Directional Reward-preserving Environment Shift Screening), a post-training framework that converts this failure mode into a statistically controlled audit. DRESS retains perturbation direction to prevent cancellation, uses paired nominal and shifted rollouts, and tests cost increase jointly with reward noninferiority. A discovery-confirmation split, hierarchical effect intervals, exact training-seed tests, and family-wise correction separate candidate selection from reproducible evidence while respecting the hierarchical structure of RL evaluation. Experiments across multiple offline safe RL algorithms, tasks, and deployment shifts reveal a previously hidden delay-induced safety-cost amplification that is invisible to a saturated binary violation metric and persists on unseen policy checkpoints. Broader stress tests and nominal-placebo analyses further demonstrate that DRESS produces selective, well-calibrated conclusions, while its adaptive acquisition strategy improves screening efficiency under limited evaluation budgets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.