GeoMapBench: Cognition-Aware Characterization of Geospatial Capabilities in Vision-Language Models
Abstract
Vision-language models (VLMs) are increasingly applied to maps, remote-sensing imagery, and other geospatial data, yet their geospatial capabilities remain difficult to evaluate systematically because existing benchmarks typically cover only subsets of cartographic understanding, spatial reasoning, geographic knowledge, and structured geospatial problem solving. We develop a comprehensive taxonomy that decomposes geospatial capability into 23 fine-grained tasks organized into four domains: Visual Perception and Interpretation, Spatial Reasoning and Geometric Processing, Geospatial Language and Semantic Reasoning, and Applied Geospatial and Earth-System Reasoning. Guided by this taxonomy, we construct GeoMapBench, a benchmark containing 2,300 examples with 100 instances per task. Each example is additionally assigned a level of the revised Bloom taxonomy, enabling analysis across both geospatial capability and cognitive complexity. We also construct GeoMapRAGCorpus, a benchmark-separated multimodal retrieval corpus containing 180,344 records from seven public geospatial source families. The corpus covers structured geographic data, text, map imagery, and geocoded geographic images. We benchmark five contemporary VLMs on all 2,300 GeoMapBench examples and observe substantial variation across tasks, with no single model consistently dominating the full task space. Finally, to demonstrate the downstream utility of the released resources, we evaluate GeoMapPlus, a task-aware retrieval- and tool-augmented inference pipeline built using GeoMapRAGCorpus. GeoMapPlus improves both task-aware and strict evaluation performance over its underlying VLM, demonstrating that the corpus can support measurable improvements in geospatial reasoning.ب
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.