Internalizing Geometric Law: Learning from Solver Residuals for Precision-Critical Generation
Abstract
Large Language Models frequently hallucinate in precision-critical domains such as technical diagramming and mechanical design, where outputs must satisfy strict geometric constraints. We study this failure where it can be checked exactly: open-ended geometric synthesis from natural language, in which an LLM must turn a free-form description into a precise geometric construction whose objects simultaneously satisfy dozens of interacting constraints. To make it trainable we release PyGeoX, a programmable geometric DSL that compiles declarative constraints into a differentiable loss, an agent harness that exposes PyGeoX to any LLM, and a 900-problem evaluation suite with per-constraint verifiable rewards spanning procedurally generated, examination and engineering problems. Using PyGeoX as a verifier, we identify a failure mode we call Outlier Gradient Masking: when residuals are aggregated through a single norm, for example , one outlier constraint can nullify the learning signal from all the others. To address this we propose Saturating Additive Rewards (SAR), which decompose the reward into bounded per-constraint terms so that progress on satisfied constraints survives. Against MSE-based rewards, the natural baseline for geometry solvers, SAR improves the hard-tier solving rate by , and the resulting 8B model is competitive with much larger frontier systems.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.