CAPO: Towards Enhancing LLM Reasoning through Generative Credit Assignment
Abstract
Reinforcement Learning with Verifiable Rewards (RLVR) has significantly improved the reasoning capabilities of large language models (LLMs). However, existing RLVR methods provide limited control over intermediate reasoning quality. Value-based approaches such as PPO suffer from inaccurate credit assignment due to limited sampling, while process reward model (PRM)-based methods rely on costly human annotations and often produce unreliable process signals. To address these limitations, we propose Credit Assignment Policy Optimization (CAPO), a simple and efficient method. CAPO avoids training auxiliary models and instead leverages an off-the-shelf, general-purpose LLM as a Generative Process Reward Model (LLM-as-GenPRM) to generate step-wise critiques in a single pass based solely on step correctness. This enables deterministic credit assignment to refine tokens that would otherwise receive identical rule-based rewards. As a result, CAPO significantly simplifies the training pipeline. % while maintaining strong generality across a wide range of powerful, publicly available open-source models. Extensive experiments on LLaMA and Qwen backbones demonstrate that CAPO consistently outperforms six strong baselines across four mathematical benchmarks and three out-of-domain benchmarks. Further analyses show that CAPO’s credit assignment mechanism substantially improves reasoning quality.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.