acceptodds
Under review as a conference paper at ICLR 2027

Hidden Sensitivity in Spatial Reasoning Benchmarks: Diagnosis and Factor-Balanced Evaluation

Abstract

Spatial reasoning benchmarks for vision-language models (VLMs) report aggregate scores that implicitly weight evaluation conditions by their frequency in the benchmark. Across four video spatial reasoning benchmarks (VSI-Bench, ReVSI, VSTI-Bench, and OSI-Bench) and 14 VLMs, we show that the distributions of task-agnostic factors, such as ground-truth answer value and queried object type, are highly skewed. Models exhibit distinct performance patterns across these factors, and plausible changes in composition alone often reorder leaderboards. We quantify this effect through a model’s composition advantage, defined as the difference between its observed score and a factor-balanced counterpart. To identify consequential factors, we introduce the Composition Sensitivity Test (CST), an information-theoretic diagnostic that retains factor-performance dependencies that persist after adjustment for entangled factors. It identifies answer value and object type as recurring sources of composition sensitivity, while also capturing additional influential factors when present. The Factor-Balanced Score (FBS) then reweights benchmark items to alleviate the dominance of frequently represented levels of diagnosed factors in the aggregate score. Across the four benchmarks, balancing changes the ranks of 53% of model-task pairs and the top-ranked model on 10 of 29 affected tasks. On object counting, for example, answers are concentrated in a narrow range of values; on VSI-Bench, balancing this distribution decreases a spatially fine-tuned model’s score by 0.25, while a general-purpose model whose score barely changes rises from eighth to first place. This illustrates how benchmark compositions can unintentionally favor some models over others. Moreover, under systematic composition changes, FBS reduces score variability by 63-71% relative to the observed scores. Together, these results show that benchmark composition can notably affect conclusions about spatial reasoning performance of VLMs. We therefore advocate jointly reporting observed scores, balanced scores, and composition advantages to make this influence explicit.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.