Less Is More: Semantic Curation of Earth Observation Data for Self-Supervised Learning
Abstract
Earth-observation (EO) foundation models are pretrained on massive satellite archives that contain substantial spatial redundancy. Yet EO pretraining data are rarely curated semantically, and are instead typically used as collected or sampled using hand-designed geographic and metadata-based rules. We investigate the impact of data curation on self-supervised EO pretraining and introduce GeoSelect, a training-free semantic curation pipeline tailored to EO. Starting from embeddings of a frozen EO foundation model, GeoSelect removes acquisition-induced variability before selecting a fixed-size subset of diverse observations. On the globally distributed Sentinel-2 Major-TOM archive, pretraining on only 10% of the data yields better representations than pretraining on the full archive. Across GEO-Bench, this improves linear-probing performance by +2.8 points in classification and +5.5 mIoU in segmentation. GeoSelect also outperforms geographic, metadata-based, and standard semantic curation baselines. These gains are robust across curation budgets, and the approach extends across sensing modalities. Our results show that semantic curation can improve EO representation quality while reducing the amount of pretraining data by an order of magnitude. We will release the curation pipeline and curated subsets.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.