Learning to Localize by Retrieval: Spectral-Spatial Contrastive Representation Learning for Multispectral Satellite Imagery
Abstract
In most cases, supervised object localization in multispectral images is based on dense bounding-box annotations which is not always possible when these annotations are not available or are costly. We formulate localization as similarity search in a learned embedding space and avoid learning to predict a bounding box, thereby eliminating the need for bounding-box supervision for detector-free localization. We train a four channel multispectral encoder with self-supervised momentum contrastive learning to learn the discriminative patch representations. A partition of the target imagery is constructed into overlapping patches, which are incorporated into a normalized representation space, and then indexed so that it is easy to retrieve the patch corresponding to the nearest imagery. Similar patches are retrieved and translated into localized regions by similarity filtering and spatial consolidation, given a query region. We test the learned representation from four complementary viewpoints: embedding separability, retrieval without labels, localization consistency despite query perturbations and retrieval scalability with the growth of the indexed corpus. Based on our results, self-supervised contrastive learning yields structured multispectral representations, which enables repeatable query retrieval based localization at larger corpses and exhibit a low query latency. The results obtained in this study show the viability of an alternative approach to localization in data-constrained multispectral remote-sensing environments based on detector-free retrieval, instead of using annotation intensive detection.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.