SkillCritic: Skill-Aware Calibration of Generative Critics for Mathematical Reasoning
Abstract
Reinforcement learning for mathematical reasoning requires finer-grained credit assignment than outcome-level rewards alone can provide. Generative critics are a promising direction because they evaluate partial solutions segment by segment, but we find that they can still assign systematically over-optimistic scores to prefixes that already contain recurring reasoning failure patterns. We propose SkillCritic, a skill-aware calibration mechanism that uses structured failure-pattern priors to convert raw segment scores into risk-calibrated token-level training signals via a bounded soft multiplicative correction, without replacing the critic or introducing an external verifier. Under a unified implementation framework, we compare SkillCritic with matched value-free, conventional critic, and generative critic baselines. SkillCritic shows its clearest gains in lightweight training, where it delivers stronger and more robust improvements, while remaining competitive under full RLPT and showing more concentrated benefits on harder contest-style math subsets. We further diagnose a characteristic instability in prolonged full-RLPT training of generative critics, and find that it is better explained by critic score collapse and downstream baseline miscalibration than by direct explosion of raw value magnitudes. Overall, our results suggest that soft skill-aware calibration is an effective inductive bias for generative-critic reinforcement learning, especially in lightweight regimes and on more challenging mathematical reasoning tasks.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.