acceptodds
Under review as a conference paper at ICLR 2027

From Map to Street: A Benchmark for Multi-Scale Geospatial Reasoning in Instruction-Driven Geographic Localization

Abstract

Instruction-driven geographic localization is a fundamental geospatial-reasoning task in which a model must fuse multiple sources of evidence to locate a real-world place described by natural-language instructions. Despite its practical importance, existing benchmarks capture only a narrow, static version of this task: they provide fixed inputs, allow little active exploration of geospatial space, and evaluate models primarily by final localization accuracy. To address this limitation, we present **IDOL**, a comprehensive benchmark for **I**nstruction-**D**riven **O**pen-world geographic **L**ocalization, which requires an agent to locate a place that a colloquial instruction describes by fusing multi-modal evidence across map, satellite, and street-level views. The benchmark covers eight cities and seven languages, and defines two task families: pinpointing a unique place from visual cues, and finding any place satisfying an open-ended predicate. For each family, we design a diverse suite of metrics that assess the reasoning process rather than the final answer alone. Experimental results show that while powerful models such as GPT-5.6-Sol and Gemini-3.1-Pro demonstrate great potential, they remain limited in fine-grained visual understanding and spatial grounding; moreover, their performance varies across cities and languages, exposing the imbalance in current models' reasoning capabilities. Our benchmark-generation and evaluation frameworks can be found at https://anonymous.4open.science/r/IDOLBenchmark/.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.