ForecastBench-Sim: Forecasting Simulated Worlds
Abstract
Predicting how a complex system will evolve is among the most general tests of reasoning under uncertainty, and among the hardest to benchmark. Real-world forecasting benchmarks often resolve slowly and provide sparse evidence of accuracy in forecasting rare or counterfactual events. We introduce ForecastBench-Sim, a forecasting benchmark that asks models to forecast the outcomes of simulated worlds. We generate questions about rare events, conditional outcomes, and responses to causal interventions. Across 24 models, three simulations, and nine distinct tasks, model scores closely track general measures of model capabilities ( []). ForecastBench-Sim provides tasks whose question sets can be refreshed, varied in difficulty, and immediately scored. We discuss potential mechanisms to prevent saturation and avoid benchmark gaming, focusing on the automatable extensibility of our framework to new unseen questions.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.