acceptodds
Under review as a conference paper at ICLR 2027

Cup Arena: Can Large Language Models Beat the Market? A Live 2026 FIFA World Cup Benchmark

Abstract

Can large language models beat betting markets in live football forecasting, and improve as match outcomes accumulate? We investigate this research question with CupArenaBench, a benchmark built from the 2026 FIFA World Cup. Across 103 matches, we require 10 frontier LLMs to submit real-time scoreline and win–draw–loss forecasts before kickoff. Each resolved match provides new evidence for subsequent forecasts. We find that most frontier LLMs outperform conventional forecasting baselines. This pattern persists when market odds are withheld. We then study whether LLM forecasters can improve as new match outcomes become available. Dynamic memory allows agents to retain lessons from resolved matches. Our Forecast Harness enables them to execute forecasting code. As matches resolve, the forecasting code evolves from the newly observed outcomes. This continual adaptation leads to progressively better forecasts as the tournament unfolds. We further introduce Football-RL, which updates model parameters from the chronological match stream. The trained policy also transfers zero-shot to the Premier League, improving over the base policy. Together, these results show that LLM forecasters can continually improve from one-shot real-world feedback as new matches resolve. We release the benchmark, data, and code.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.