acceptodds
Under review as a conference paper at ICLR 2027

Primal-only Safe Policy Improvement at Test-time

Abstract

Safe reinforcement learning typically synthesizes a cost constraint into policy parameters during training, and the deployed policy does not re-check the constraint at the state being visited. We ask whether the constraint can instead be imposed at test time on sampled candidate actions. An action constructed at decision time can be checked against the current cost critic at the visited state, without waiting for a policy or a dual variable to update. We propose safe test-time extraction of policies (STEP), which draws candidates from a behavior prior and refines each with a primal-only switch. Specifically, a candidate follows the reward gradient while its estimated cost is below the threshold and receives a cost correction otherwise. The gate evaluates each candidate as a single deviation followed by the behavior policy, and safety under repeated STEP execution is assessed empirically. On DSRL tasks, the average cost of STEP over seeds stays within the budget on , and it attains the highest return among agents within the budget on . In online fine-tuning on five tasks, STEP maintains or slightly improves return while reducing cost, without an initial performance drop.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.