AτA: A Benchmark for Multi-Agent interaction in Agentic Economies
Abstract
Agents will increasingly take on economically impactful, long-horizon tasks, but most current benchmarks evaluate them in isolated interactions with a single user. While such evaluations are useful for measuring whether agents can accurately accomplish tasks, they fail to capture the broader impact of deploying these agents in real-world marketplaces. In real markets, task completion alone is not enough. Outcomes depend on how well agents satisfy the preferences and constraints of multiple entities with competing objectives. In this work, we introduce Aτ A, a benchmark for multi-agent interaction in agentic economies, in which both users and providers navigate markets through autonomous agents. We evaluate user and provider utility in both symmetric (the same model represents users and providers) and asymmetric (different models represent the two sides) settings, studying how model choice, resource availability, and coordination affect the resulting Pareto trade-offs. Our results show that high booking rates do not necessarily translate into good outcomes for both sides. Even strong models leave substantial room to improve joint outcomes, revealing limitations that evaluations focused on task completion alone do not capture
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.