ZAPD: Zero-Shot Adaptation via Primal-Dual Updates for Constrained Reinforcement Learning
Abstract
In constrained reinforcement learning (CRL), agents maximize expected reward while ensuring the expected cost does not exceed a given threshold. Most existing CRL methods, including the widely used primal-dual approaches, assume this threshold is fixed, but in real-world applications it can shift abruptly. Standard CRL methods then need substantial re-learning, and existing threshold-adaptive methods require access to a range of thresholds or a library of pre-trained policies. To address this problem, we propose Zero-shot Adaptation via Primal-Dual updates (ZAPD), which adapts the primal-dual solution to a new threshold without any additional environment interaction. ZAPD treats the shift as a perturbation of the converged solution and employs a first-order sensitivity analysis to derive closed-form primal and dual updates. Both updates are computed from data the underlying method already collects, so ZAPD plugs into a broad range of primal-dual CRL algorithms. Experiments show that ZAPD achieves the lowest immediate approximation error on all four Bullet Safety Gym tasks in the single-threshold setting, and that this improved initialization translates into faster post-shift convergence than threshold-adaptive baselines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.