Veri-GPS: Learning to Coordinate Geometric Proof and Calculation with Verification and Solver Feedback
Abstract
Large language models (LLMs) can solve geometric proof problems with the support of symbolic solvers. However, many geometry problems require more than correct geometric relations, as proof and calculation must also support each other. We introduce Veri-GPS, an interaction learning framework that learns subgoal policies from symbolic solver feedback to coordinate proof and calculation within a given interaction budget. An extended and accelerated Newclid backend enables geometric proof and exact calculation to share verified results. When execution is blocked, goal-related diagnostic feedback identifies missing premises and quantities along candidate rule or formula paths to limit LLM trial and error. To train the LLM to use solver feedback, we construct a dual-source pipeline combining teacher–solver interaction trajectories with Goal Augmentation trajectories. Both sources are split into turn-level training data for supervised fine-tuning (SFT) of Qwen2.5-VL-7B-Instruct, providing a cold start for the interaction policy. We introduce turn-level Dynamic-Difficulty Reinforcement Learning (DDRL), combining adaptive sampling across fixed difficulty buckets with derivation-use rewards for intermediate results used in subsequent reasoning. Across 3,082 problems from Geometry3K, PGPS9K, the planar geometry subset of MathVerse, GeoQA+, and GeoQA, Veri-GPS achieves a problem-weighted pass@1 of 82.12%. Compared with standalone symbolic execution and Qwen3-VL-32B-Instruct direct answering, Veri-GPS improves performance by 10.25 and 17.39 percentage points, respectively. Within Veri-GPS, full feedback yields a 1.27-point gain over reduced feedback at inference with the trained policy fixed, while DDRL improves pass@1 by 3.57 points over the SFT checkpoint.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.