Reasoning Selectively over Forecast Errors with Evidence-Conditioned Probabilistic Correction
Abstract
A probabilistic forecast can be well calibrated on average yet remain unreliable under a particular structural state. We propose evidence-guided selective probabilistic reasoning, a framework that separates the choice of a residual condition from the estimation and deployment of its numerical correction. A numerical forecaster provides a predictive center and asymmetric interval. A selective multi-agent workflow examines bounded structural evidence and produces a residual ConditionSpec, which deterministic estimators use to retrieve mature residuals and compute center and tail corrections. An outcome-grounded Gate then selects how much of the candidate to deploy using comparable historical candidates frozen before their outcomes were known. Delayed outcomes populate separate views for residual estimation, workflow allocation, and deployment. Experiments on XuanCheng and PEMS-BAY show dataset-dependent effects of evidence conditioning and an explicit accuracy-computation trade-off. Selective reasoning reduces live LLM calls by 38.7% and 60.2% relative to Always FULL, respectively, with higher MAE than Always FULL but lower point-prediction errors than the Base. On PEMS-BAY, suppressed candidates have negative mean interval-score gain, whereas fully deployed candidates have positive mean gain. These results connect structural reasoning to measurable candidate quality and deployment behavior, providing a traceable way to control both reasoning effort and forecast correction.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.