acceptodds
Under review as a conference paper at ICLR 2027

Low-Loss Temporal Connectivity in Weather Foundation Models

Abstract

AI weather models are usually trained for one fixed forecast step: a 6-hour model outputs only 6-hour forecasts. To obtain a forecast at an in-between horizon such as 4 hours, one currently has to retrain a new model, chain several short-step models, or interpolate their outputs on the grid. Each option is costly, blurs fine spatial detail, or both. This is a real obstacle for downstream systems (ocean, sea-ice, hydrology and air-quality models, and sub-6-hour data assimilation) that need forcing at cadences unrelated to 6 hours. We show empirically that intermediate horizons can instead be recovered directly in a model’s weights. Starting from a public 6-hour global model checkpoint (Aurora), we fine-tune 3-hour and 1-hour versions with an identical recipe. Comparing internal representations with Centered Kernel Alignment (CKA), the 3-hour and 6-hour models are nearly identical, consistent with a shared low-loss region of weight space, while the 1-hour model has moved to a different region. Along the straight line between the 3-hour and 6-hour weights, a single number λselects a forecaster whose effective lead time tracks λ. We verify this at the two interior horizons that hourly ERA5 offers on this segment: at 0.25◦, plain linear weight interpolation (LWI) at 4 and 5 hours beats every grid-space output-interpolation baseline on near-surface variables, with up to about 2×less error energy at mesoscale wavelengths. Against autoregressive chains the trade-off is the expected one: the chain has lower geopotential point error but damps the spectrum, while LWI preserves the observed energy spectrum at one-third of the chain’s inference cost. Spherical-interpolation and Fisher-weighted variants give no improvement. Connectivity is local: it holds on the 3–6 h segment, decays as the horizon gap grows, and fails under extrapolation. The effect reproduces on two further, architecturally different models (Pangu-Weather and GraphCast) and is stable across training seeds. We present this as an empirical characterization of when weight-space interpolation across lead times works and where it fails; its use as forcing for coupled systems or in assimilation is the motivation, not a demonstrated result.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.