acceptodds
Under review as a conference paper at ICLR 2027

Anchored Latent Predictions for Earth Foundation Models

Abstract

Self-supervised Earth-observation encoders learn scene context, but mapping clouds and floods also requires local detail. In two runs of a latent-prediction recipe with weak pixel-target weights, extending training from 10B to 40B tokens cost 5.6 and 5.7 points of frozen-probe cloud mIoU; thin-cloud IoU fell by 28 to 34%, and crop-type and flood segmentation also degraded. Meanwhile, land-cover and GEO-Bench accuracy peaked at the worst cloud checkpoints, and feature rank rose. Better scene recognition can therefore hide a loss of spatial detail, motivating fixed pixel-level targets that keep evolving representations tied to the observations. We propose anchored latent prediction (ALP): complement evolving latent targets with strongly weighted, frozen random projections of the pixels, predicted at twice the token resolution with no extra encoder cost. In single-run tests, switching to ALP’s loss weights, including a higher cloud-loss weight, halted the decline; an ALP run kept improving until about 20B tokens. EarthFM, a 91M-parameter en- coder trained with ALP plus map and cloud-model targets, is on par with or better than OlmoEarth v1.2 Base on six of eight public test sets from GEO-Bench, GEO-Bench-2, and Sen1Floods11, with 1.8 to 3 times less hardware-adjusted pretraining compute: ahead on clouds, within noise on five tasks, and behind on EuroSAT and So2Sat. Its shorter 18.75B-token branch leads on clouds with a third of OlmoEarth’s compute (+1.15 mIoU, 95% CI 0.53 to 1.79); only TerraMind-L scores higher, with 7.2 times as much compute.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.