acceptodds
Under review as a conference paper at ICLR 2027

GeoRefer-Bench: A Benchmark from Referring Pixels to Verifiable Geospatial Reasoning

Abstract

We introduce GeoRefer-Bench, a benchmark for evaluating whether vision-language models identify the objects denoted by spatial expressions in remote sensing imagery. Pixel overlap measures segmentation quality but does not directly assess whether the predicted object set satisfies a spatial query. GeoRefer-Bench links natural-language expressions to executable logical forms over annotated scene graphs, providing explicit query semantics and traceable answer sets. It comprises 1,710 scenes from georeferenced UAV imagery and three re-instrumented public datasets, with 50,488 objects and 49,512 executable queries covering five query structures, from entity selection to two-hop relations. Unanswerable queries, counterfactual pairs and paraphrases enable targeted tests of spatial reference resolution. We complement pixel overlap with Exact Query Success (EQS), assessing object selection through a shared mask-to-object read-out that uses ground-truth instance masks at evaluation time. Relation-blind baselines expose a discrepancy between overlap and correct selection: on answerable UAV relation queries, predicting all instances of the named category achieves 36.9 mIoU but only 9.2% EQS. Evaluation of fifteen models on 1,881 UAV test queries reveals substantial performance gaps between entity and compositional queries: the strongest overall model reaches 98.9 at entity selection but only 60.5 at two-hop composition. Stratified analyses also associate greater same-category ambiguity with lower success, even within answer-area strata. Together, these results motivate evaluating spatial referring segmentation through both pixel accuracy and query-consistent object selection.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.