Web2Earth: Extending Web Vision Foundation Models Across Remote Sensing Modalities
Abstract
Vision foundation models pretrained on large-scale web imagery learn rich visual representations, yet their RGB-centric training leaves a gap in their ability to handle other sensing modalities. Motivated by the pursuit of remote sensing foundation models that perform well across modalities and tasks, we ask whether this existing knowledge can be extended beyond RGB while also improving optical performance. We introduce Web2Earth, a framework that builds on a web-pretrained encoder to pursue broad remote sensing utility through joint RGB self-distillation and cross-modal knowledge transfer. We design complementary patch and relational distillation objectives to transfer local spatial knowledge and scene-level relationships from RGB to other sensing modalities. Adaptation uses 32k RGB images and approximately 5k co-registered image pairs, while updating only approximately 2% of the model parameters. Evaluations across diverse optical and SAR tasks demonstrate strong cross-modal transfer, with leading results on the main optical evaluations and the highest average score across six SAR evaluations among the compared models. Notably, joint RGB adaptation and cross-modal alignment improve optical performance over the starting web-pretrained model while extending its utility to SAR. These results support extending the knowledge in large-scale web-pretrained models toward remote sensing foundation models with strong performance across both modalities and tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.