acceptodds
Under review as a conference paper at ICLR 2027

Marginal Scoring Hides What Generative Forecasters Learn

Abstract

Probabilistic multivariate time-series forecasting is dominated by a family of models that pair a deterministic backbone with a generative component, and that report CRPS and QICE as headline numbers. We show that both metrics, as implemented throughout this literature, are computed independently at every (time step, channel) element: they score the predictive marginals and are blind, by construction, to dependence across the forecast horizon. We make the consequence concrete. Taking a recent state-of-the-art model and holding its backbone, normalisation, training objective, sampling budget and evaluation code fixed, we replace only the generative sampler with a diagonal Gaussian head that cannot represent any temporal dependence. This deliberately dependence-free floor beats the published model on CRPS in of datasethorizon cells (by –, at seed standard deviations of at most ) and is generally better calibrated; the four exceptions are the two highest-dimensional benchmarks, where its deficit is –, and a lookback sweep shows that which side wins CRPS is a regime question, decided jointly by dimensionality and context length. Under a dependence-sensitive diagnostic there is no regime: the lag-1 autocorrelation of the sampler's residual paths — what it adds beyond the conditional mean, which is zero for the floor by construction — is positive for the generative model in every one of the cells we ran, across seven benchmarks, two horizons and three lookbacks. It is also small: the true forecast-error autocorrelation is – everywhere, and the generative model recovers – of it on the - and -series benchmarks, and all of it only on the two high-dimensional ones — exactly the cells where it also wins CRPS. The two metric families therefore rank the same pair of models in opposite directions, and the literature reports only one of them. We then isolate the cause. Across three datasets, giving the sampler an explicit horizon covariance buys – less autocorrelation than changing the objective does, so the objective, not the architecture, is the binding constraint. Replacing it with a joint proper scoring rule recovers – of the ground-truth autocorrelation (versus – for the published model), and an energy-score/KDE hybrid dominates the published model on CRPS, QICE, energy score and variogram score simultaneously on all three.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.