BazaarBench: Delegation Safety in Decentralized C2C Marketplaces Run by LLM Agents
Abstract
In decentralized consumer-to-consumer (C2C) marketplaces, people list goods, negotiate with strangers, and rate one another, so trust rests on reputation. Large language model agents now act and decide for users. On a real platform, their unsafe choices would cost users money, privacy, and reputation. Even among humans, unsafe trades are common and hard to spot, since weak policing lets people fake shipments or talk buyers into scams. The key question is whether agents trading for people fail in such unsafe ways. BazaarBench, a simulated C2C marketplace, keeps each item's owner, true condition, and commitments apart from agents' claims. It links six failure types to trades, so the truth of every trade can be observed. We first create 100 synthetic personas with inventories from an open inventory dataset, and three models each run one base market of these agents for 30 simulated days. Then 20 of the 100 agents in each market switch to one of five tested models, keeping personas, inventories, and histories. They trade for seven days under ordinary instructions, deadline pressure, or a user's adversarial instructions, each in its own market copy. Unsafe trading is routine even without harmful instructions, as 16% to 22% of committed deals in the base markets go through despite a failure the platform's records confirm. No tested model is safe by default. In the ordinary baseline after the switch, where models of different capability trade together, each causes such failures in its own deals, 14% overall and 8% to 23% by model. Deadline pressure pushes them into riskier purchases, most often of units the seller has also promised to another buyer. Under adversarial instructions, the share of tested sellers' committed deals carrying a fake or wrong item to the handoff more than doubles, from 15% to 33%. Had inspection not stopped them, their weekly gain would have risen by about 60%. Inspection catches almost all of them, so completion falls and realized gain stays flat, while the few that deceive a buyer are shipped without inspection. We therefore recommend checking ownership and commitments at acceptance and requiring inspection or delivery evidence before completion. The 48 runs hold 357,608 agent model calls with complete market records, and code and data are available for analysis and future training at https://anonymous.4open.science/r/BazaarBench.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.