acceptodds
Under review as a conference paper at ICLR 2027

Training Mathematical Proof Generators via Reference Free Generative Verifiers

Abstract

Mathematical reasoning goes beyond finding correct answers to constructing rigorous proofs. Recent efforts have used reinforcement learning with verifier feedback to improve language models' ability to generate such proofs. However, the long reasoning traces required for generation and verification make reinforcement learning in this setting difficult to stabilize. Through ablations on smaller models, we identify three choices that improve training stability: sequence-level rather than token-level loss aggregation, large batches, and adaptive positive-advantage masking. This masking method controls policy entropy by selectively suppressing updates to low-probability tokens with positive advantage. Starting from Nemotron 3 Ultra, we train a proof generator using feedback from a fixed DeepSeek-Math-V2 verifier that receives only the problem and candidate proof. The resulting model surpasses DeepSeek-Math-V2 itself in proof-generation accuracy on downstream evaluation. The gains persist under an independent judge and reference-aware evaluation.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.