Overcoming Skill Credit Blindness: Skill-Conditioned Policy Optimization for LLM Reasoning
Abstract
Learning reusable reasoning skills through reinforcement learning (RL) is essential for improving both training efficiency and the ability of large language models (LLMs) to solve challenging reasoning tasks. However, most existing methods rely on Group Relative Policy Optimization (GRPO), whose group-relative baseline can absorb systematic skill effects, obscuring whether performance gains arise from policy improvement or injected skills. We term this credit-assignment failure skill credit blindness. To address it, we propose Skill-Conditioned Policy Optimization (SCPO), which explicitly conditions policy optimization on skills and reconstructs the advantage signal for reliable attribution. To isolate policy quality from skill effects, SCPO normalizes rewards only within the same problem and skill condition. It then uses paired skill-conditioned and skill-free rollouts to estimate each skill’s instance-level marginal contribution, which modulates the advantage to reinforce beneficial skills and suppress harmful ones. Finally, SCPO computes cross-problem skill-level gains to support the evolution of the skill library. Experiments with Qwen3-series models across multiple mathematical reasoning benchmarks show that SCPO consistently outperforms GRPO, demonstrating skill-based RL depends not only on learning reusable skills, but also on correctly attributing skill contributions during policy optimization.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.