acceptodds
Under review as a conference paper at ICLR 2027

Beyond Static Criteria: Model–Rubric Co-Evolution for Reinforcement Learning with Self-Distillation

Abstract

Recently, on-policy self-distillation (OPSD) and its reinforcement-learning extensions have offered a promising route to improving large language models. By conditioning the teacher on additional privileged information, these methods provide token-level guidance along the student's own sampled trajectories. However, the effectiveness of these methods depends on predefined privileged teacher information, giving rise to two key challenges. First, teacher contexts are often constructed independently of each sampled rollout, making them unable to adapt to its specific errors or unmet behaviors. Second, the teacher information is typically fixed before training and cannot evolve with the improving capabilities of the policy. To address these limitations, we propose **M**odel–**R**ubric Co-Evolution for **R**einforcement **L**earning with **S**elf-**D**istillation (MR.RLSD), an integrated reinforcement-learning and self-distillation framework centered on co-evolving rubrics. MR.RLSD uses rubrics as a shared interface between reward feedback and teacher guidance. Rubric scores determine policy-update directions, while criteria unmet by each response are transformed online into teacher contexts that provide contrastive feedback for modulating token-level update magnitudes, without requiring pre-constructed privileged contexts. During training, a shared policy serves as both a response generator and a rubric generator, proposing new criteria based on current rollouts and historical rubrics. The updated criteria jointly refine reward computation and teacher guidance, allowing both the optimization targets and self-distillation feedback to adapt to the evolving policy. Experiments across health, writing, and science show that MR.RLSD outperforms the evaluated reinforcement learning and self-distillation baselines, demonstrating the effectiveness of co-evolving the policy and its guiding criteria for open-ended generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.