CriPO: Enhancing Rubric-based RL via Self-Distillation
Abstract
Rubric-based Reinforcement Learning (RL) has recently shown promise in improving Large Language Models (LLMs) on open-ended tasks. A widely recognized limitation of rubric-based RL is limited exploration: criteria that no rollout manages to satisfy (*Unexplored Criteria*) receive no optimization signal. Recent methods address this by incorporating rubric information as external guidance during rollout generation, yet they introduce a *train-inference mismatch*: the policy is optimized on rollouts produced under external guidance while this guidance is absent at inference time, causing error accumulation through autoregressive decoding. Moreover, these exploration-focused approaches overlook a fundamentally different failure mode that we term *Suppressed Criteria*—criteria that are satisfied by some rollouts yet whose learning signals might be lost during optimization because scalar reward aggregation assigns them non-positive aggregate advantages. Our analysis reveals that suppressed criteria constitute a persistent and non-negligible failure mode—over 25% of samples exhibit this issue throughout training. To simultaneously address both unexplored and suppressed criteria without introducing training-inference mismatch, we propose **Criterion-Distilled Policy Optimization (CriPO)**, which enhances rubric-based RL via on-policy self-distillation. For unexplored criteria, CriPO constructs a behavior-injection teacher and computes a filtered forward-KL loss to inject missing behaviors into the policy. For suppressed criteria, CriPO uses a counterfactual teacher to locate criterion-relevant tokens in negative-advantage rollouts, and corrects their advantages in GRPO to preserve useful patterns. Experiments on medicine and science benchmarks demonstrate that CriPO outperforms existing rubric-based RL methods, e.g., achieving an average gain of **3.3** points over GRPO on Qwen3-4B.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.