When Agents Spend for Users: Benchmarking Economic Decision Security in A2A Marketplaces
Abstract
Large language model (LLM) agents increasingly choose among paid services offered by other agents, making provider selection economically consequential. Existing benchmarks primarily evaluate task completion and rarely test whether Host Agents follow economic decision rules, enforce user policies, or remain reliable when Providers strategically present information under their control. We introduce A2AEcoBench, a benchmark for economic decision security in Agent-to-Agent marketplaces. It evaluates Normative Decision Competence—economic rationality, user-policy compliance, and trusted-source alignment—and Adversarial Manipulation Robustness, defined as reference-consistent selection under controlled Provider-side interventions involving names, promotional descriptions, and direct instructions. A2AEcoBench contains 20,000 controlled provider-selection instances spanning 10 task scenarios and 20 subtasks. Each instance follows a pre-declared reference decision rule, enabling reproducible evaluation without assuming a universal economic utility function. On A2AEcoBench, eight Host Models obtain a mean Overall Score of 49.6/100. User Policy Compliance has the highest raw accuracy at 88.7, whereas accuracy under Promotional Description conditions is lowest at 32.8. Decision-intensive scenarios score 10.9 points below utility-oriented scenarios on average. These results show that task completion and explicit policy following do not by themselves guarantee reliable economic delegation, and motivate provenance-aware selection, enforceable decision authority, and defenses against economically motivated manipulation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.