acceptodds
Under review as a conference paper at ICLR 2027

MapSpatial: A Benchmark for Representation-Robust Spatial Reasoning over Real Maps

Abstract

Reliable map use requires stable answers across views and orientations, and revised answers when roads or buildings change; single-image accuracy cannot measure both. We introduce MapSpatial, a geographic information system (GIS)-grounded benchmark of 77,800 instances spanning 14 sub-tasks in four families: qualitative direction, metric comparison, route-relative building identification and counting, and road-network path reasoning. Because every scene is stored as GIS geometry, each answer can be recomputed exactly when the same question is posed across three registered views (satellite imagery, a text-free structural map, and a conventional text-bearing map), under rotations and mirror flips, and, where applicable, after solver-verified structural edits paired with answer-preserving shams. We evaluate 25 multimodal models. Among six representative models, direct accuracy is 35.2–48.3%, but 33.8–60.7% of registered-view triplets yield different semantic predictions across views that share an answer but differ in visual cues. After a single declared geometric action, 21.1–34.3% of base-correct answers become wrong; the stricter base-plus-seven joint score is 11.2–22.8%. After an answer-changing edit, 9.6–26.2% of base-correct pairs update correctly, while 50.4–84.3% retain the pre-edit answer. Gemini-3-Flash correctly updates 50.4% of its base-correct edit pairs. Controlled visual aids show task-dependent effects. MapSpatial quantifies cross-view disagreement, transform sensitivity, and limited correct updating through paired, GIS-verifiable views and edits.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.