Primal-only Safe Policy Improvement at Test-time
Abstract
Safe reinforcement learning typically synthesizes a cost constraint into policy parameters during training, and the deployed policy does not re-check the constraint at the state being visited. We ask whether the constraint can instead be imposed at test time on sampled candidate actions. An action constructed at decision time can be checked against the current cost critic at the visited state, without waiting for a policy or a dual variable to update. We propose safe test-time extraction of policies (STEP), which draws candidates from a behavior prior and refines each with a primal-only switch. Specifically, a candidate follows the reward gradient while its estimated cost is below the threshold and receives a cost correction otherwise. The gate evaluates each candidate as a single deviation followed by the behavior policy, and safety under repeated STEP execution is assessed empirically. On DSRL tasks, the average cost of STEP over seeds stays within the budget on , and it attains the highest return among agents within the budget on . In online fine-tuning on five tasks, STEP maintains or slightly improves return while reducing cost, without an initial performance drop.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.