acceptodds
Under review as a conference paper at ICLR 2027

Learning to Escalate: Do LLM Agents Generalize Governance Policy to Unseen Rules?

Abstract

LLM agents that act through tools are usually held to a written policy by a runtime gate that blocks rule-breaking calls. A gate covers only the rules someone has encoded, so we ask what an agent learns about a policy from reward alone, and whether that learning reaches rules it was never trained on. We train Qwen2.5-7B-Instruct with GRPO on customer-support tasks from τ²-bench's telecom domain, rewarding task success and penalizing policy violations and escalations to a human. To test transfer, we remove entire rule families from every channel through which training could expose them. The held-out rule we study requires the customer's consent before a paid data top-up, and the agent learns only part of it. It learns restraint, making no unauthorized top-ups on any of three seeds. It does not learn the handoff the policy asks for; instead, training on other rules erodes a handoff it had learned during fine-tuning, from 99% of blocked requests to 8%, and a single-seed test points to the escalation cost as a major cause. It also becomes reluctant to top up even when allowed. An agent trained on every rule family avoids all three failures and matches the gate on safety and task success. The held-out agent commits no violations, so a violation rate would call it compliant. We argue that policy evaluations should also check whether the agent hands off what it may not do.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.