SIGHT-Bench: Diagnosing Reference-Frame Consistency and Global Spatial Reasoning in Multimodal Models
Abstract
We introduce SIGHT-Bench, a diagnostic benchmark for spatial reasoning in real-world elevated outdoor scenes. It draws on stationary camera installations whose positions remain fixed while viewing orientations may vary. SIGHT-Bench contains 1,000 four-choice questions from 400 cameras, organized into five dimensions and 12 subtasks spanning geometry, scene layout, viewpoint changes, world-frame motion, and topology. We evaluate proprietary, general-purpose open, and spatially specialized multimodal models against a human reference. Even the strongest evaluated models retain substantial overall gaps to human performance, and those gaps vary across spatial dimensions. Across model configurations, spatial consistency is more reliable than direct connectivity. Additional thinking yields marked gains in world-frame motion but much smaller gains elsewhere. Subtask analyses further show that improvements in one spatial relation need not extend to related judgments. These results motivate treating spatial reasoning as a profile of distinct capabilities rather than a single aggregate score.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.