acceptodds
Under review as a conference paper at ICLR 2027

DeFuture: Decoupling LLM Forecasting with Hindsight Evidence

Abstract

Future event prediction tests whether a large language model (LLM) can turn information about the present into a judgment about the future. Current benchmarks score each system end-to-end, so a failed prediction cannot be traced to specific components, and an agentic forecaster cannot tell whether to improve its search, its processing, or its reasoning. We introduce **DeFuture**, which decouples future event prediction into evidence acquisition, interpretation, and prediction, and anchors this decomposition on golden evidence that a selector picks in hindsight from the same answer-blind evidence pool by reading the resolved answer. Golden evidence makes every phase measurable: it tops a ladder of four evidence tiers, carries known answer-bearing content for scoring each Interpreter, and gives all Predictors a shared input. On 1,254 FutureX-Past questions and 340 MIRAI relation queries, DeFuture shows that (1) decoupling measures each phase at almost no cost in accuracy, (2) evidence weighs more than model capability, since a 4B Interpreter and a 27B Predictor match a frontier model given golden evidence, and (3) even frontier Predictors miss over 25% of questions whose golden evidence contains the answer, with extra reasoning helping only at prediction. We release the golden evidence, evidence pools, model outputs, and code.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.