FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents
Abstract
Language model agents now execute bounded tasks reliably, but whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, is less well studied. We introduce FM-Bench (Football Management Benchmark) to study this setting. An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops, drafting on a fixed budget, trading, negotiating contracts, investing in facilities and youth, and answering to a board that can fire it. A deterministic engine accumulates every season into one final score with no LLM judge. A solo track plays 15 frontier models against a frozen scripted world; an Arena places them in one shared 20-year world. Across three solo seeds, all models complete the horizon, whereas the blind scripted baselines are eliminated in most runs; claude-fable-5 has the highest mean solo score and finishes first in the Arena, although the season title rotates among ten models. In our sample, neither model scale, price, nor vendor clearly predicts the ranking, which stabilizes only late in the horizon, and the best of six first-time human players scores near the bottom of the model board. Decomposing the score into six behavioral capabilities suggests that differences in managerial behavior, rather than computation, account for much of the gap: stronger models tend to reduce slow-payoff investment near the end, keep cash invested, and renew contracts early. Token spend is not correlated with score, no model reliably infers the market's hidden prices, and self-managed memory degrades into an ever-growing archive or a plan rewritten every season. The code is available at https://anonymous.4open.science/r/FM-Bench-B644.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.