acceptodds
Under review as a conference paper at ICLR 2027

SpatialGraph-Eval: Intervention-Based Evaluation of Spatial-Network Reasoning in Multimodal Models

Abstract

Multimodal language models increasingly perform well on spatial and visual reasoning benchmarks, but aggregate accuracy does not show how model performance changes when the same problem is presented differently or when its underlying structure is modified. We introduce SpatialGraph-Eval, a diagnostic benchmark for spatial-network reasoning that separates graph structure from its presentation and enables targeted interventions on both. Graphs are procedurally generated with exact shortest-path solutions and controlled variation in path length and decision complexity, while graph size and mean branching are matched within path-length conditions. We evaluate three contemporary multimodal model systems on identical graph structures presented symbolically and visually, under topology-preserving rotations and reflections, and following minimal topology-changing counterfactuals. Representation sensitivity differed substantially across models, with optimal-route accuracy decreasing by 3.9 percentage points for Terra (68.1% to 64.2%), 42.8 for Claude (85.6% to 42.8%), and 17.2 for Gemini (100.0% to 82.8%) when comparing visual with symbolic presentation. Despite unchanged topology, optimal-route accuracy varied substantially across rotations and reflections, with significant heterogeneity across models (Model × Transform, p = 1.83 × 10⁻¹¹). Among instances solved optimally before intervention, models found the new shortest route in only 2.0–7.1% of cases after a one-edge counterfactual rewiring changed the correct route. Failure analysis showed that degradation included graph-invalid transitions as well as valid but suboptimal routes. As an external validity check, we evaluated 50 subgraphs sampled from the San Francisco OpenStreetMap road network; across models, optimal-route accuracy decreased from 90–100% under symbolic presentation to 2–16% under visual presentation. These results reveal substantial model-dependent behavioral sensitivity to representation, topology-preserving presentation changes, and minimal structural interventions, motivating intervention-based evaluation of spatial-network reasoning beyond aggregate task accuracy.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.