AlphaBet: Reinforcement Learning for LLM Event Forecasting Agents
Abstract
Forecasting an unresolved event requires finding relevant evidence and deciding how strongly it supports each outcome. Yet outcome-based training often starts from preassembled evidence at a single prediction date, leaving evidence acquisition outside the learning loop. We introduce **AlphaBet**, a reinforcement learning framework that trains a shared LLM policy to search, reason, and report probabilities across multiple pre-resolution dates. Each dated forecast receives proper-scoring feedback from the eventual outcome, allowing one resolved event to supervise decisions under different information states. We decompose forecasting regret into gaps from evidence acquisition and probability reporting, and show that GRPO standard-deviation normalization can distort the reporting target under multi-class proper scores. On GLM-4.5-Air, **AlphaBet** reduces Brier from 0.731 to 0.626 on Forecast-Dojo and from 0.612 to 0.528 on FutureX. Controlled ablations attribute these gains to longitudinal training and improvements in both acquisition and reporting, while showing that GRPO standard-deviation normalization harms probabilistic forecasting while accuracy remains comparable.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.