GEBench: Benchmarking Tool-Augmented Agents for Geo-Economic Reasoning
Abstract
abstract Geo-economic decision-making, like public-facility assessment and allocation, requires not only interpreting multimodal geospatial information but also actively discovering what data is missing and which geographic tools can supply it. While recent geospatial foundation models and multimodal large language models have improved perception and geo-grounding, the existing benchmarks predominantly focus on single-modal inputs and tasks related to remote sensing, leaving tool-driven economic and demographic knowledge discovery underexplored. This paper therefore introduces the Geo-Economic Benchmark (GEBench), a large-scale multimodal benchmark for predictive geo-economic reasoning grounded in remote sensing imagery, map-derived place context, and demographic priors. To the best of our knowledge, this is the first work to benchmark agentic, tool-interactive methods for solving geo-economic problems. GEBench formulates geo-economic tasks as question-answer pairs with multi-modal input and supports multi-round interaction. It comprises over 16,000 question-answer pairs across five building categories with their corresponding geo-economic prediction targets and spanning more than 6 countries. Building on GEBench, we develop an agentic reasoning pipeline in which an exploration agent plans feasible API queries and iteratively integrates results from mapping and demographic tools before downstream analysis and forecasting, enabling active information acquisition from Geo-Economic data. We systematically evaluate tasks including numerical prediction and localization, demonstrating that our agentic reasoning framework consistently outperforms direct model querying across diverse geo-economic scenarios.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.