Latent Weather Forecasting with Joint-Embedding Predictive Architectures
Abstract
We investigate Joint-Embedding Predictive Architectures (JEPAs), a self-supervised approach that predicts in embedding space rather than pixel space, as a basis for data-driven weather prediction. We train JEPA-WX on 39 years of ERA5 reanalysis at 5.6° resolution for 12 upper-air variables at three pressure levels. The architecture combines a Vision-Transformer context encoder, an exponential-moving-average (EMA) target encoder, a Transformer predictor with free-running rollout training, and a lightweight convolutional decoder. Training uses three terms: a latent prediction loss, a Slicing Univariate Test regularizer against representation collapse, and an auxiliary reconstruction loss that we find essential for downstream decoding. We show that (i) JEPA-WX outperforms persistence at all evaluated lead times on all 12 variables, with a day-1 anomaly correlation (ACC) of 0.94 on Z500; (ii) JEPA-WX is competitive with a pixel-space U-Net at short leads and outperforms it at day 14 on 8 of 12 variables; and (iii) RMSE grows by 3.2× from day 1 to day 14 for JEPA-WX versus 7.6× for the U-Net, suggesting that the latent bottleneck acts as an implicit regularizer against autoregressive error accumulation. Code, checkpoints and evaluation scripts are released.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.