Do AI Weather Models Miss Extremes? Evidence from Ten Months of Verification Against European Weather Stations
Abstract
AI weather models are often reported to underpredict extremes, but most evidence comes from deterministic regression systems verified against reanalysis. We compare twelve physical and AI forecast systems with ECMWF IFS using European station observations over ten months. The evaluation covers 10 m wind, 2 m temperature, solar radiation, and precipitation, with regimes defined from a fixed ERA5 1991–2020 climatology and models compared on matched samples. We find no uniform AI-specific loss of relative skill in the tails. Tail performance varies substantially among both learned and numerical systems: several AI systems remain more accurate than IFS in extreme regimes, while others deteriorate markedly. This conclusion persists across variables, although rankings are sensitive to forecast range, regime definition, and mean-bias correction. All systems also show a similar observation-conditioned error structure, overpredicting low observations and underpredicting high observations. Extreme-regime performance is therefore specific to the forecast system and evaluation setting rather than a uniform weakness of AI weather forecasting.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.