Improving Long-Horizon Forecasts with Non-Markovian Corrections to Pretrained Weather Models
Abstract
Unresolved atmospheric processes can introduce history dependence beyond the short input windows used by many weather forecasting models. We introduce GraphCast–Mamba, which augments a frozen, pretrained \(1^\circ\) GraphCast backbone with a trainable Mamba correction branch combining graph message passing and Mamba layers. The branch retains a temporal memory state across successive six-hour forecast steps, learning corrections informed by temporal context beyond GraphCast's two-frame input. Training is confined to the correction branch, using ERA5 data from 2015–2021. After only 2,500 optimizer updates, including warm-up, the selected model achieves substantial improvements relative to the uncorrected GraphCast backbone. It reduces the GraphCast normalized weighted mean-squared error (MSE; GraphCast loss) by 16.3% when aggregated over the full ten-day forecast and by 23.7% at day 10 (240 h after forecast initialization). At this final forecast lead time, root-mean-square error (RMSE) decreases by 10.2–17.1% across all 11 evaluated variable groups, spanning temperature, wind, pressure, geopotential, specific humidity, and precipitation. Separate memory-reset controls support a contribution from persistent temporal context beyond additional trainable capacity. These results demonstrate broad improvements in ten-day weather forecasts through a short correction-training schedule while retaining the pretrained backbone.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.