Delegation-Scoped Provenance for Tool-Using Agents
Abstract
Users often let an agent take the arguments of an action from what it reads: "set up the meeting in Quinn's email", "message everyone in #general", "refund whoever paid me for the deposit". Defenses against indirect prompt injection either forbid untrusted content from choosing such arguments, which blocks these tasks, or judge whether an action looks legitimate. We instead read the request as a statement of delegated authority: it says, often implicitly, which records and lookups may supply which arguments. DelegationScope parses this delegation from the request alone, admits a sensitive value only if a record the user authorized contains it, and asks a verifier that sees only authorized records to settle values that other content also mentions. We describe attacks that defeat record-level provenance checks, and DelegationBench, 216 tasks whose delegation is known. In preregistered held-out evaluations, DelegationScope stops all 96 attacks that plant the attacker's value in the named record, 42 of which pass a provenance baseline, and on AgentDyn it lowers attack success from 138 to 81 of 560 while completing more tasks than that baseline, though fewer than no defense. Across five executors on AgentDojo, it cuts attack success by 29–52% where injections succeed and costs at most three of 97 clean tasks where they do not. DelegationBench shows where reading delegation fails; the revised DelegationScope+ lowers attack success to 9 of 194 on AgentDojo, from 31 without defense and with no loss of clean tasks, and to 62 on AgentDyn. CaMeL and Progent, run in the same setting, block more attacks with a 9B executor but complete 6 to 69 fewer of the 97 tasks, mostly ones whose arguments the user delegated. Injections inside the record the user named remain beyond any provenance rule.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.