FantasyClash: Benchmarking Strategic Reasoning in Fantasy Sports
Abstract
An intelligent agent that operates among its peers must reason strategically: it must model how its counterparts will respond to its own action. Existing evaluations of strategic reasoning in LLM agents rely on stylized matrix games, board and card games, or dialogue-based negotiation, which lack the complex dynamics that govern the uncertainty real-world agents tackle. We present FantasyClash, a benchmark for strategic reasoning under real-world uncertainty. The benchmark is built on the popular fantasy sports format, where agents (1) bid for athletes in an auction, (2) field weekly lineups from their pools, and (3) bet on their teams' performance relative to others. Crucially, the outcomes are determined by the corresponding athletes' performances in real games, which imbues the game with high-dimensional, real-world dynamics. Evaluating 8 frontier LLMs, we find that agents vary substantially in their innate ability to play the game. We further write detailed descriptions of 7 strategies and verify them as effective prompts for improving an agent's performance. We empirically identify intransitive cycles among a subset of strategy-prompted agents, illustrating strategic depth where success depends on inferring counterparts' strategies. Furthermore, we find that agents rarely act on what their opponents do, but some agents exploit an opponent's strategy once they see its full text. Finally, we characterize the relative strengths of the 8 agents through a large-scale evaluation with an Elo rating system.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.