acceptodds
Under review as a conference paper at ICLR 2027

SketchForge: Prover-Guided Reinforcement Learning for Sketch Generation in Agentic Theorem Proving

Abstract

Sketch generation is a key capability of LLM-based theorem-proving agents. However, training sketch generation models remains hindered by unreliable correctness assessments, rewards misaligned with difficulty reduction, and limited lemma-level provability feedback for sketch refinement. To this end, we introduce SketchForge, an agentic reinforcement learning method for sketch generation. SketchForge comprises three components: (1) a prover-guided sketch verification framework; (2) sketch effectiveness discrimination through prover selection; and (3) a model-as-a-tool system for learning iterative sketch refinement from lemma-level feedback. Trained with our method, SketchForge-27B improves Pass@32 over its base model by an average of 4.7 percentage points across five benchmarks, matching or surpassing OProver-32B, the previous SOTA open-weight agentic prover of comparable size, with substantially less training data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.