acceptodds
Under review as a conference paper at ICLR 2027

What Should We Forecast? Benchmarking Agents on Early Question Discovery

Abstract

Forecasting benchmarks usually assume that the question has already been written. We study the upstream problem: can an agent discover which questions should be asked before the relevant events become obvious? We introduce AnteBench, a benchmark that evaluates whether an agent formulates forecasting questions in advance about the uncertainties that later lead to consequential developments. Using only news published before a cutoff date, the agent produces a list of forecasting questions about the next 30 days; after the 30-day window closes, we construct a reference list of questions from the available information and match the two lists. We use average precision (AP) to evaluate how much of the reference list the agent covers and how highly it ranks those questions in its own list. In a 30-day backtest covering Eastern Europe and Russia in April 2026, we evaluate frontier agents and find that performance differs across agents and improves as reasoning effort increases. We also show that AnteBench serves as a testbed for experimentally studying agents' abilities to gather information and formulate relevant forecasting questions from it.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.