acceptodds
Under review as a conference paper at ICLR 2027

Taxonomy-Driven Map Query Agent Design and Evaluation

Abstract

Large language models (LLMs) equipped with tools can resolve map queries expressed in natural language. However, existing systems often assemble their tools and evaluation around target benchmarks, inheriting their capability gaps and lacking an independent reference for assessing coverage. We address this limitation with a faceted taxonomy of capabilities for map queries over 2D vector map data, validated through LLM and human-expert annotation on three datasets. From this taxonomy, we systematically construct a twelve-tool inventory with minimized functional overlap and synthesize evaluation data that covers every capability it defines. We build a simple agent to compose these tools into executable query plans. Experiments across three LLM families show that our agent improves on baseline results for two external benchmarks while maintaining relatively stable accuracy across planner scales. The ablation shows that overlapping tool interfaces reduce accuracy more for smaller planners, supporting the taxonomy's role in the systematic construction of tool inventories with minimal functional overlap.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.