RAFMa: Retrieval-Augmented Flow Matching for Global Visual Geolocation
Abstract
Global visual geolocation predicts where on Earth an image was taken. Classification and generative models store all world knowledge in their parameters, cannot be updated without retraining, and are known to generalize well at coarse scales, such as country or region. In contrast, retrieval-based systems compare the image against a database of geotagged images and excel at fine scales, such as street or building, where memorization matters; moreover, their database can be updated at any time as new geotagged images become available. We combine the two families in a flow matching model conditioned jointly on the query image and on retrieved visually similar images together with their geotags, a form of in-context learning. Starting from a random location on the sphere, the model iteratively moves it along a learned velocity field, while an attention module learns to weight each retrieved image according to its appearance and geodesic distance to the current estimate. The model achieves a good trade-off between generalization and memorization, and effectively benefits from knowledge updates at inference time by changing the database. On street-view images (OSV-5M), RAFMa sets a new state of the art, clearly outperforming prior work. On unconstrained user photographs (YFCC4k, Im2GPS3k), it is competitive with systems built on multimodal LLMs of several billion parameters, while training only 135M parameters.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.