acceptodds
Under review as a conference paper at ICLR 2027

Policy-Verified Self-Training: Learning Agent Compliance from Verifiable Policy Rewards

Abstract

Runtime policy verifiers can prevent agents from executing disallowed actions, but their decisions are typically used only for enforcement and discarded afterward. We ask whether these decisions can instead become verifiable policy rewards for training the agents they guard. We introduce Policy-Verified Self-Training (PVST), a post-training paradigm in which an agent generates its own interaction trajectories and an executable policy verifier, here a deterministic gate compiled from the natural-language policy, deterministically evaluates whether each action complies. These verdicts provide verifiable training signals: compliant trajectories are reinforced, and violations are used to construct negatives for preference optimization. PVST thus turns an existing runtime guardrail into a teacher for the same agent: beyond the gate and the benchmark’s task-success signal, it needs no stronger teacher model, no human preference labels and no attack-specific supervision, although rollouts are collected in the benchmark’s attacked environment. We evaluate PVST on tool-using agents under prompt injection. On heldout AgentDojo tasks, policy-verified preference training of Qwen3-8B reduces violating trajectories from 37.2% to 10.3 ± 0.9% (three seeds) and executed attacks from 12.4% to 0.5% at equal or higher task utility, compared with 22.0% violating trajectories when negatives are selected by an attack oracle and 31.9 ± 0.9% with matched negatives that contain no violation (a success-only control that never consults the gate lands at 20.2%). This separation indicates that the improvement comes from the policy information encoded by the verifier rather than from preference optimization alone. The learned behavior also transfers beyond the training distribution: it holds under an attack template never used in training (37.7% to 11.0%), violations decrease from 34.8% to 16.7% on a suite whose rule set and tasks were excluded from training, and measurable improvements emerge from a few dozen preference pairs. When the original verifier is retained at deployment, its intervention rate falls from 39.2% to 8.8%, reducing reliance on runtime enforcement while preserving a deterministic safety backstop; the compliance result replicates on a second trainee (Qwen3-4B-Instruct-2507, 37.3% to 10.2%). Our analysis identifies the condition under which the verdict is most informative: guarded actions must appear in both permitted and forbidden contexts. Executable policies can therefore serve not only as runtime constraints but as a scalable source of verifiable rewards for self-training, and they motivate learning objectives that exploit the verifier’s finer-grained rule- and predicate-level structure.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.