BridgeGeom: Diagnosing Fine-Grained Geometric Perception under Structural Aliasing
Abstract
Vision-language models often recognize a bridge as suspension, cable-stayed, or arch, but this does not show that they perceive the geometry that identifies the bridge itself. We study this gap under structural aliasing, where objects share a coarse engineering template and identity rests on weak local geometric cues. We introduce BridgeGeom, a diagnostic benchmark of 971 bridges, 7,555 standardized images, and 10,826 questions, whose core tasks fix the structural type so that only geometric evidence can distinguish instances. Scaling from Qwen3.5-4B to 27B lifts name-image assignment from 19.00% to 32.30%, but same-bridge verification rises only from 70.90% to 75.90% and height comparison from 55.40% to 60.30%. A structural equation over the task measurements yields a dominant scale factor that follows a scaling law with model size, and the geometric blind spot hides inside this factor. We separate it by decomposing the factor into general, image-text, and geometric capability, and measure the geometric capability at 1.057 for Qwen3.5-4B, 1.281 for 27B, and 1.396 for GPT-5.6-sol. Retrieval augmentation lifts the image-text capability but never the geometric capability, and the blind spot is thus contained within the scaling law yet separable from overall capability and shortcut use.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.