SolidGeoGen: Advancing Multimodal Solid-Geometry and Spatial Reasoning through Verifiable Symbolic Generation
Abstract
Solid geometry requires multimodal large language models (MLLMs) to jointly interpret 3D structures and perform mathematical and spatial reasoning, yet high-quality data with aligned diagrams and verifiable solutions remains scarce. Existing data-generation methods focus primarily on plane geometry, while solid-geometry problem generation remains limited in both diversity and reliability. We present SolidGeoGen, a theorem-driven symbolic framework for generating diverse solid-geometry problems with aligned diagrams, formal representations, and verifiable reasoning traces. This design ensures high-quality, internally consistent data for both training and evaluation. Building on SolidGeoGen, we construct SBench, a Solid-Spatial Benchmark comprising 2,000 problems across ten solid-geometry and spatial-reasoning tasks. We further build STrain, a training corpus containing 10,000 examples. Leveraging STrain, we propose VISTA, a Visual-structure-Informed and Symbolic-Trace-Assisted training framework. Extensive evaluation highlights the difficulty of SBench: all evaluated open-source MLLMs achieve below 30% average accuracy, while even the strongest proprietary model reaches only 67.2%. VISTA substantially improves accuracy on SBench at both 9B and 27B scales, while achieving consistent gains on external benchmarks for multimodal math reasoning and spatial intelligence. These results demonstrate that symbolic generation can provide scalable and reliable supervision for advancing complex multimodal reasoning in MLLMs.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.