TradingArena: A Market-Making Benchmark for Ranking, Diagnosing, and Steering LLM Behavior
Abstract
As large language models are increasingly deployed as agents, evaluation should capture both task performance and interactive behavior. Few benchmarks jointly measure performance and behavior under controlled changes to an agent’s information, action space, and inference. We introduce **TradingArena**, a multi-agent market-making environment for both performance ranking and controlled behavioral analysis. Agents trade on a shared limit order book using noisy, biased observations of a changing hidden value, with performance measured by mechanically computed profit and loss (PnL). We evaluated 14 models across six configurations and multiple competitive rosters for performance ranking, and further used controlled interventions to diagnose behavioral failure modes. In particular, we quantitatively diagnose over-confidence in GPT-5.5, which trades aggressively under unreliable information and shows a low abstention rate on unanswerable questions. A self-refine intervention reduces this aggressiveness and moves GPT-5.5 from last to second among 14 models by mean PnL in a matched evaluation. These results illustrate how TradingArena can reveal not only differences in model outcomes, but also how behavior changes under intervention.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.