acceptodds
Under review as a conference paper at ICLR 2027

Decodable Is Not Grounded: A Grounding Arbiter for VLM Spatial Reasoning

Abstract

Spatial representations in vision–language models (VLMs) can support effective interventions while the resulting answer gains persist without the image. Across fourteen VLMs on ViewSpatial-Bench, thirteen score below chance on depth and twelve perform better on horizontal relations than depth. Matched-protocol interventions in three VLMs reveal a shared reversed depth response: steering toward the positive relation pole lowers its answer probability, while reverse-sign oracle injection improves accuracy by 13.7–22.2 percentage points. Vertical direction removal instead has model-dependent effects. We introduce the Grounding Arbiter to examine what successful edits recover: compare matched, blank, and mismatched images, then repeat the same intervention with matched and blank inputs. In a detailed case study, vertical relations reach 86.9% probe accuracy under scene-grouped nested evaluation, yet an amplification operating point selected from an exploratory sweep gains 20.5 points with real images and 19.1 with blank images. We call this an ablation-persistent intervention gain. Further analyses show that decoding and controllability can peak at different layers, and cross-model task comparisons reveal large image benefits on other depth benchmarks. These findings connect the evaluation of spatial representations to two practical choices: where to intervene, and how to measure the resulting gain's dependence on visual evidence.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.