CARTA: A Traceable Benchmark for Cross-Representation Map Understanding in Multimodal LLMs
Abstract
The same place can be shown as satellite imagery, a street map, a land-cover map, a contour sheet, or a weather chart. Reading across these visual languages is a basic requirement for models that navigate, monitor, or respond in the physical world, yet existing benchmarks rarely change the representation while holding the geographic referent fixed. We introduce CARTA, a benchmark that asks what breaks when the place stays the same but its representation changes. CARTA derives images, questions, answers, and distractors from shared geospatial state, so each of its 16,408 items is traceable from source data through rendering, audit, and frozen image bytes. The items form matched comparison families for representation profiling, continuous-field reading, and cross-representation correspondence, together with breadth and grounding tracks. Evaluating twelve MLLMs against a matched human reference reveals failures at three successive stages. Evidence: human accuracy on forest and cropland/bare targets falls to 15–45% on street maps, and an objective purity screen shrinks the human satellite–minimalist gap from 20.0 to 8.6 pp. Correspondence: models match categories more readily than places, as cross-class distractors inflate correspondence accuracy by up to 22 pp (Qwen3-VL-8B: 34.7% to 56.9%), and every model loses 3.5–14.4 pp when the two views are not aligned. Geometry: on point grounding, open-weight models do no better than copying the source mark's coordinates (17.6% versus 17.5%), and even GPT-5.5 reaches only 28.9%, against 88.3% for readers on a matched subset. On matched multiple-choice items GPT-5.5 scores 74.4% against the readers' 83.3%. Reasoning-enabled serving lifts two open-weight models from about 61% to 71% on MCQs but provides no reliable benefit for geometric grounding. CARTA thus provides a matched test bed for distinguishing representation-specific success from transferable geographic understanding. Data and code will be released after acceptance.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.