Can language models forecast inductively?
Abstract
Judgmental forecasting involves predicting uncertain events like elections or geopolitical conflicts where abundant historical data is lacking. A key challenge is induction: forecasters must generalize from sparse evidence and contextual knowledge to infer how available evidence predicts future outcomes. Here, we investigate whether language models (LMs) recover and use such predictive relationships when forecasting, rather than relying on base rates or surface cues. Across a battery of experiments spanning real-world prediction markets and controlled synthetic environments, we test whether models: 1) update their beliefs coherently; 2) correctly recover predictive relationships, and 3) transfer what they learn to new settings. Our results provide evidence for each capability. Models selectively update their forecasts in response to causally relevant evidence, use contextual information to infer predictive relationships when observations are sparse, and revise these inferences as direct evidence accumulates. Training on families of predictive relationships improves forecasting across new domains, although evidence for transfer to unseen predictive structures is weaker. Together, our experiments provide behavioral evidence that LMs can forecast inductively.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.