acceptodds
Under review as a conference paper at ICLR 2027

UAV-MapVQA: Evaluating Foundation Models on Instance-to-Entity Grounding Question Answering across UAV Imagery and Vector Maps

Abstract

Current aerial perception benchmarks largely evaluate class-level recognition within the image plane, such as detection, counting, and scene understanding, but do not test whether models can ground visual instances to their corresponding real-world entities in a vector map. This capability is critical for UAV applications that rely on vector maps as compact, structured world knowledge for linking observations to semantics, attributes, and spatial context. We introduce UAV-MapVQA, the first benchmark for cross-view instance-level grounding and geospatial reasoning between UAV images and vector maps. The benchmark contains 336K multiple-choice QA pairs generated from 14,000 real-world UAV images collected across 3 countries (USA, China, and Germany) and more than 1,000 km of flights, using an automatic registration and formally verified QA generation pipeline with human validation. It spans eight tasks across four levels, from cross-view instance–entity grounding and change detection to viewpoint-aware localization and object-centric geospatial reasoning. Each question is provided with two alternative representations of the same vector map—a rendered raster map and a structured entity table—enabling controlled study of how map representation affects cross-view grounding and geospatial reasoning. Evaluating 13 foundation models (FMs), spanning proprietary and open-weight general-purpose models such as GPT-5.5 and Claude Opus 4.7 as well as domain-specific geospatial FM GeoChat, we find that no zero-shot model exceeds 50% overall accuracy, against 94.8% for humans. Supervised fine-tuning on UAV-MapVQA raises Qwen2.5-VL-32B from 23.0% to 57.2%, surpassing all zero-shot models, which confirms that the dataset provides an effective training signal; yet a large gap to humans remains, and aligning the map to the UAV heading shows that orientation errors propagate into downstream reasoning but explain only part of this gap. These findings expose major gaps in cross-view instance grounding and map-conditioned geospatial reasoning, establishing UAV-MapVQA as a new evaluation and training resource for map-aware multimodal intelligence at the intersection of aerial robotics, multimodal learning, and geospatial AI. An interactive demo of the benchmark is available at an anonymous URL: https://storage.googleapis.com/uav_map_vqa/index.html.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.