Beyond Generation and Accuracy: Diagnosing and Enhancing Visual Chain-of-Thought for Geometry Problem Solving
Abstract
Solving geometric problems often requires more than passively reading a given diagram; constructing auxiliary lines is crucial to reveal hidden geometric relations. While Visual Chain-of-Thought (VCoT) allows models to introduce such intermediate visual aids, a key question remains: why do autonomous systems struggle to benefit from them? In this work, we present the first systematic study that decouples the potential value of visual aids from a model’s ability to construct and effectively use them. Specifically, we evaluate models under three controlled setups—No-Aux, Auto-Aux, and GT-Aux—along with a five-dimensional trajectory diagnosis on our proposed benchmark, GeoVAD-Bench. Our analysis reveals a critical bottleneck in VCoT, which we call the autonomy gap: while high-quality auxiliary diagrams consistently boost accuracy, autonomous models fail to capture these gains when generating aids on their own. Diagnostic results show that errors compound across perception, construction, utilization, and deduction, which ex- plains why visual quality alone cannot account for this performance gap. Driven by these insights, we introduce a targeted post-training paradigm and the first reinforcement learning algorithm designed for interleaved visual-textual reasoning. Our model, GeoWeave-8B, outperforms the base model by 25.3 percentage points in final geometric accuracy and 30.4 percentage points in process diagnostic metrics. These results demonstrate that a diagnosis-guided approach provides a clear path toward closing the autonomy gap in geometric reasoning. Code is available in the supplementary material.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.