Training Mathematical Proof Generators via Reference Free Generative Verifiers
Abstract
Mathematical reasoning goes beyond finding correct answers to constructing rigorous proofs. Recent efforts have used reinforcement learning with verifier feedback to improve language models' ability to generate such proofs. However, the long reasoning traces required for generation and verification make reinforcement learning in this setting difficult to stabilize. Through ablations on smaller models, we identify three choices that improve training stability: sequence-level rather than token-level loss aggregation, large batches, and adaptive positive-advantage masking. This masking method controls policy entropy by selectively suppressing updates to low-probability tokens with positive advantage. Starting from Nemotron 3 Ultra, we train a proof generator using feedback from a fixed DeepSeek-Math-V2 verifier that receives only the problem and candidate proof. The resulting model surpasses DeepSeek-Math-V2 itself in proof-generation accuracy on downstream evaluation. The gains persist under an independent judge and reference-aware evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.