SketchForge: Prover-Guided Reinforcement Learning for Sketch Generation in Agentic Theorem Proving
Abstract
Sketch generation is a key capability of LLM-based theorem-proving agents. However, training sketch generation models remains hindered by unreliable correctness assessments, rewards misaligned with difficulty reduction, and limited lemma-level provability feedback for sketch refinement. To this end, we introduce SketchForge, an agentic reinforcement learning method for sketch generation. SketchForge comprises three components: (1) a prover-guided sketch verification framework; (2) sketch effectiveness discrimination through prover selection; and (3) a model-as-a-tool system for learning iterative sketch refinement from lemma-level feedback. Trained with our method, SketchForge-27B improves Pass@32 over its base model by an average of 4.7 percentage points across five benchmarks, matching or surpassing OProver-32B, the previous SOTA open-weight agentic prover of comparable size, with substantially less training data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.