How much ensemble does a frozen weather model contain?
Abstract
Data-driven weather models now match numerical weather prediction, and to represent uncertainty, the field trains probabilistic models end to end, at the full training cost of the forecaster. We ask instead how much ensemble skill a pretrained deterministic model provides with its weights frozen, measured against its own forecast. HyperENS trains low-rank adapters on the backbone’s attention projections and a hypernetwork that generates their cores from a Gaussian latent: each new latent is a new member, whose adapter stays fixed along its forecast. The member distribution is trained end to end with an almost-fair CRPS on the forecasts and their power spectra, and no spread amplitude is tuned afterwards. Applied to the pretrained 0.25° Aurora, HyperENS+, which also adapts the encoder’s query latents, trains 3,235,544 parameters, 0.26% of the model, within two days on eight A100 GPUs, and leaves the original deterministic forecast unchanged. Scored against ERA5 on 178 initialisations of 2024, averaging per-variable relative scores over eight variables, its 32-member ensemble mean has 18.8% and 24.7% lower RMSE than the frozen forecast at days five and ten, and its fair CRPS is 0.569 and 0.525 times that forecast’s mean absolute error. Aurora 1.5’s ensemble, a whole-model fine-tune, has the lower absolute fair CRPS: ours is 9.6% and 6.8% higher at days five and ten. The mean spread–skill ratio is within six percent of one from day three, but short leads remain under-dispersed (0.861 at day one) and the rank histograms show residual errors. On an earlier checkpoint evaluated on 418 common initialisations, the adapter’s static part alone is only slightly better than the frozen forecast; most of the gain appears only in the ensemble.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.