Disagreement-Adaptive Safe Policy Updating under Protected Subgroup Constraints
Abstract
Safe policy improvement from offline data seeks to improve upon an incumbent policy while satisfying prespecified safety constraints. Existing work has studied high-confidence policy improvement under multiple fairness or safety constraints and disagreement-aware policy comparison, while it remains unclear whether candidate–incumbent disagreement can sharpen finite-sample certification across protected subgroups. To address this gap, we study safe policy updating when candidate policies may differ from the incumbent only on a small fraction of each evaluation population. We show that direct estimation of candidate–incumbent value differences yields simultaneous confidence bounds whose leading width scales as , where is the disagreement probability and is the audit sample size, and establish a protected-subgroup lower bound together with a local minimax result showing that this dependence cannot be uniformly improved near a binding subgroup constraint. We then develop DiSPU, a diagreement-adaptive framework for safe policy updating (DiSPU) that incorporates these confidence bounds into a constrained deployment optimization, and prove high-probability safety and disagreement-dependent oracle-regret guarantees. Extensive experimental results confirm that exploiting policy disagreement reduces certification conservatism and enables larger certified improvements while maintaining protected-subgroup safety.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.