PTPO: Progressive Theorem-Guided Policy Optimization for LLM-Based Mathematical Reasoning
Abstract
Despite recent great advances in mathematical reasoning, large language models still struggle with competition-level problems that require selecting appropriate mathematical theorems and applying them under the correct conditions. Existing reinforcement learning methods mainly optimize final-answer correctness or generic reasoning trajectories, often providing limited guidance on the theorem-level principles. To address this limitation, we propose Progressive Theorem-guided Policy Optimization (PTPO), a theorem-guided reinforcement learning framework with input-adaptive progressive prompt construction. We first construct a structured theorem library and generates theorem-grounded chain-of-thought prompts that explicitly connect problems with relevant mathematical principles. Based on the theorem library, we then adaptively inject theorem-grounded prompts according to the question difficulty and model's current ability, offering rich hints for hard questions and progressively reducing hints as the model improves. This progressive schedule encourages the model to transition from external theorem guidance to autonomous theorem application. Extensive experiments on nine benchmarks show that our proposed PTPO achieves the best overall performance, and it enables faster convergence and stronger reasoning. Our source code and theorem library are available at https://anonymous.4open.science/r/PTPO-FED5.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.