Cup Arena: Can Large Language Models Beat the Market? A Live 2026 FIFA World Cup Benchmark
Abstract
Can large language models beat betting markets in live football forecasting, and improve as match outcomes accumulate? We investigate this research question with CupArenaBench, a benchmark built from the 2026 FIFA World Cup. Across 103 matches, we require 10 frontier LLMs to submit real-time scoreline and win–draw–loss forecasts before kickoff. Each resolved match provides new evidence for subsequent forecasts. We find that most frontier LLMs outperform conventional forecasting baselines. This pattern persists when market odds are withheld. We then study whether LLM forecasters can improve as new match outcomes become available. Dynamic memory allows agents to retain lessons from resolved matches. Our Forecast Harness enables them to execute forecasting code. As matches resolve, the forecasting code evolves from the newly observed outcomes. This continual adaptation leads to progressively better forecasts as the tournament unfolds. We further introduce Football-RL, which updates model parameters from the chronological match stream. The trained policy also transfers zero-shot to the Premier League, improving over the base policy. Together, these results show that LLM forecasters can continually improve from one-shot real-world feedback as new matches resolve. We release the benchmark, data, and code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.