acceptodds
Under review as a conference paper at ICLR 2027

RePolicy: Learning to Select and Apply Safety Policies for Agent Safeguards

Abstract

Agent safeguards must determine whether an agent’s behavior violates the safety rules of the current task, environment, and user. Because these rules vary across contexts, a safeguard needs to select the relevant policy before judging the behavior. Existing policy-aware safeguards can apply a supplied policy, but generally assume that the correct policy is already known. They therefore do not learn how to select a policy from a library, even though an incorrect choice can lead to an incorrect safety judgment. We introduce RePolicy, a safeguard that jointly learns policy selection and safety assessment. Given an agent trajectory and a set of candidate policies, RePolicy selects the policy that best matches the situation, applies its requirements, and predicts whether the behavior is safe. We construct PolicyTraj-20K, a dataset of more than 20,000 agent trajectories annotated with applicable policies, safety labels, and explanations. We train RePolicy with supervised demonstrations followed by reinforcement learning, while randomizing policy order, identifiers, and distractors to reduce shortcut learning. Across six agent-safety benchmarks and 19 baselines, RePolicy-4B achieves leading overall safety-detection performance and improves policy-selection accuracy. These results show that safeguards can learn which safety requirements to use for assessing agent behavior. Code is available at https://anonymous.4open.science/r/RePolicy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.