REAP: Self-Improving Policies for Attacker-Benefiting Resource Capture in LLM Agents
Abstract
Large language model (LLM) agents are expected to use user-authorized resources to complete tasks, yet existing evaluations primarily assess task success, overlooking whether the resource allocation aligns with user intent. However, even when tasks are completed correctly, attackers can redirect attention and traffic toward their own artifacts or induce unnecessary service calls, which can induce exposure or revenue for attackers while task-success metrics fail to detect this resource misuse. We study attacker-benefiting resource capture, in which adversaries redirect user-authorized resources toward attacker-controlled artifacts or services while retaining task utility. We consider two roles: (i) individual attackers redirect resources toward their own artifacts; and (ii) provider attackers increase billable use of their services. We introduce REAP, a policy-level recursive self-improvement framework that uses task execution traces to refine the policy under an explicit utility constraint and resource capture goal. REAP achieves the highest resource capture among the evaluated methods across three environments and two attacker roles, with utility preservation ranging from 90.00% to 100.00%, a citation share of 55.00%, and service-call increases of up to 1,210.00%. In our tool-use evaluation, the tested defenses expose role-dependent trade-offs, and no fixed resource budget provides a stable safety-utility balance, demonstrating task utility alone is insufficient to identify execution-time resource risk.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.