PolicyFork: Task Success Can Hide Policy Violations in Agent Delegation
Abstract
AI agents use protocols such as Agent2Agent (A2A) to delegate tasks across independently built systems. Task policies govern which agents may use a tool or service and under what conditions. Runtime systems can check submitted operations against these policies. Before that check, the receiving agent chooses an operation using a handoff, the task and context sent with the delegation. If the handoff omits a required record, the agent may reject permitted work or attempt a forbidden operation. Agent evaluation must distinguish this information gap from a model error. To address this problem, we present PolicyFork, the first framework to verify whether agent handoffs determine policy decisions. PolicyFork adds the fewest available fields needed to determine each candidate action's decision and rechecks the repaired handoff. We also construct decision-sufficient execution state (DSES), a shortest fixed-length encoding that communicates these decisions to the recipient. We implement eight software-delivery policies from GitHub, GitLab, Amazon Web Services, Azure, and Google Cloud in local A2A workflows. Our release tests show that repaired handoffs let a rule-based receiver accept 28/28 permitted state–action combinations, up from 16/28, with zero false Permits. In paired trials, adding records raises correct Permits on allowed requests from 0/48 to 48/48 for each of four model configurations. PolicyFork separates missing information from model errors and exposes policy violations that task-success scores miss. Our data and code will be made available at https://anonymous.4open.science/r/PolicyFork-ICLR27/
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.