CoRE-TSF: When to Trust Time-Series Forecasts via Cross-Modal Consistency
Abstract
Aggregate forecast accuracy does not reveal whether a particular forecast can be trusted. Yet black-box point forecasters may expose only their numerical predictions, providing no direct signal of which individual forecasts are likely to fail. We introduce CoRE-TSF, an output-level framework that cross-checks each numerical forecast against an LLM’s qualitative expectation of future change. Given the same observed history, a textual pathway predicts the direction and magnitude of future change from categorical history descriptors, while a numerical pathway forecasts future values. Neither pathway observes the other’s output. CoRE-TSF maps both outputs to a shared ordinal space and uses their disagreement to compute a forecast-level risk score without requiring realized future values, historical forecast-error labels, or access to forecaster internals. Our primary method, CoRE-Ensemble, averages risk scores from four LLM agents before ranking forecasts. We evaluate five time-series foundation forecasters across 22 dataset families, with two LLM-based forecasters evaluated on a separate matched short-horizon cohort. On the primary cohort, CoRE-Ensemble achieves a hierarchically macro-averaged within-configuration Spearman correlation of and normalized selective utility of , where measures the fraction of improvement over random retention attainable under ideal error-based ordering. These results support cross-modal consistency as an externally observable signal for ranking forecast reliability.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.