Forecast Quality Does Not Certify an Action Simulator
Abstract
Forecasting models are increasingly used to predict demand under hypothetical actions, effectively treating them as action simulators. Yet accurate factual forecasts do not guarantee reliable simulation. We identify two blind spots. First, per-step forecast evaluation can leave cross-step dependence unconstrained, even though it shapes uncertainty in trajectory-level quantities such as total demand. Second, evaluation under observed actions does not constrain how a model responds when those actions change. On two public retail panels, standard multivariate proper scores largely reproduce marginal-accuracy rankings. A matched rank-zero ablation improves weighted quantile loss (WQL) from to , while coverage of total demand falls from to and the detected dependence signal disappears. Fine-tuning improves factual forecasts for Chronos-2 and Moirai at every confounding level, but their intervention estimates differ: Chronos-2 shifts from improvement under randomized actions to degradation under confounding, whereas Moirai improves throughout. Decoder rankings also reverse under a second action-response mechanism, showing that simulator comparisons depend on the response being simulated. These findings motivate separate evaluation of marginal accuracy, cross-step dependence, aggregate calibration, and interventional fidelity rather than treating forecast quality as a certificate of simulation quality.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.