Hedge-Bench: Can Agents Reason Their Way to Better Predictions?
Abstract
Modern AI agents are increasingly capable of mechanical financial-analysis tasks such as retrieving documents and updating spreadsheets. A harder challenge is forecasting an unknowable future from incomplete information. Public markets provide a natural testbed: analysts routinely predict discrete events from point-in- time information. We present Hedge-Bench, a benchmark of 92 real-world tasks derived from hedge fund analysts’ reasoning traces and paired with subsequently resolved prediction questions. Agents reason from frozen information snapshots, and their forecasts are evaluated against realized outcomes. Across 11 frontier models and three trials per task, joint success rates range from 2.54 to 31.52%, requiring both a correct forecast and over 60% research coverage. Using GRPO with turn-level credit assignment, we train Muse-Glimmer on 165 environments and increase coverage on 22 held-out tasks from 17.17% to 33.36% at low reasoning effort and from 30.19% to 47.02% at extra-high effort. The trained model at low effort also exceeds the base model’s extra-high-effort coverage while using 47.9% fewer accepted trajectory tokens. We also observe evidence of generalization: held-out forecasting accuracy at high effort improves by 10.06 percentage points, while external evaluations improve by 3.94 points on GDPval, 2.29 on GDP.pdf, 7.07 on Harvey LAB, and 4.96 on Frontier Finance under their respective metrics.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.