Seeing in 2D, Reasoning in 3D: What Explicit Geometry Reveals about Vision-Language Models
Abstract
Does making geometry explicit remove the need for vision and is it enough for 3D reasoning? End-to-end accuracy alone cannot distinguish the benefit of supplying geometric information from the ability to solve problems once that information is available. We introduce GeoText, a structured representation for diagnostic input interventions over four geometric components: view elements, 3D shape structure, given measurements and constraints, and spatial relations. Using OrthoMind-Bench, a collection of 5,357 Chinese K–12 solid-geometry problems, we evaluate six open-weight vision-language models under single-component and cumulative interventions, each with and without images. Providing 3D shape descriptions yields the largest standalone gains across all six models, improving accuracy by 13.2–28.7 percentage points; substantial gains persist even after view descriptions are supplied. Additional givens provide much smaller gains once views and shape are available. Explicit geometry also changes the marginal value of vision: images improve accuracy by 10.5–27.9 points without GeoText, but their contribution becomes near-zero or negative in four models when all GeoText components are supplied. Yet, with both images and all GeoText components, model error rates remain 32.6–47.8%. These results reveal a separation between reduced visual dependence and successful geometric problem solving: external geometric descriptions can substantially reduce the benefit of images while leaving considerable failures unresolved. GeoText provides a diagnostic framework for measuring which geometric information improves performance, how it changes the value of vision, and what failures persist under explicit geometric support.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.