acceptodds
Under review as a conference paper at ICLR 2027

One Relation Is Not Enough: Benchmarking Relational Spatial Understanding in Vision-Language Models

Abstract

Spatial understanding requires more than recognizing isolated spatial relations; it also requires understanding how diverse descriptions involving different reference objects and reference frames jointly characterize the same spatial structure. A correct answer under a single formulation is therefore insufficient to reflect a model's ability to identify multiple valid spatial descriptions completely and accurately. We introduce CoSpatial-Bench, a direct multi-select benchmark that evaluates whether vision-language models can select all valid descriptions of the same target from a fixed candidate set, based on the scene and question, while rejecting unsupported options. The Base variant contains 495 questions and 2,000 human-reviewed ordinary candidate descriptions. Our primary metric, Error-Free Coverage (EFC), measures coverage of valid descriptions under a no-false-selection constraint, while paired variants examine recognition performance under changes in referring expressions, viewpoints, and spatial conditions. Across ten VLM baselines, the best Base EM is only 33.54, while the highest EFC is 42.53. A separate single-candidate intervention study characterizes correction of initially incorrect option decisions under perceptual support and subsequent combined perceptual and reasoning support: across three diagnostic cohorts, perceptual support corrects 58.7–70.7% of initial errors, while 8.6–13.8% remain unresolved after both stages. We further introduce Target-Centered Relational QA Synthesis (TCR-QAS) and train Qwen3.5-4B with whole-question GRPO on the synthesized data, increasing its Base EFC from 28.17 to 33.37. These results highlight the need to evaluate spatial understanding beyond the correctness of a single answer: complete and error-free identification of diverse valid descriptions remains challenging.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.