acceptodds
Under review as a conference paper at ICLR 2027

CoPE: Co-Evolving Policy and Experience via Context Distillation for Open-Ended Generation

Abstract

Reinforcement learning with rubric-based rewards is widely used for open-ended generation tasks. However, scalar rewards discard the rich information contained in language feedback. Context distillation aims to internalize knowledge from external context, which we refer to as experience, into model parameters. Fixed, pre-defined context can become suboptimal as the policy improves during training. We propose CoPE, an on-policy context-distillation framework that jointly improves a single policy model's downstream-task performance and rubric-generation capability. The evolving rubrics are then used by a frozen feedback model to extract adaptive experience for subsequent policy training. CoPE includes two phases of context distillation: a) training the policy model with experience as additional context and b) training the policy model as the rubric generator to produce better rubrics for experience extraction. The two phases are performed iteratively and enhance each other during training. We evaluate CoPE on open-ended generation tasks and evaluation ability, where the policy needs to generate rubrics to choose the preferred response. Our results show that CoPE outperforms scalar reward RL and previous experience learning approaches and co-evolves the policy model on downstream tasks and rubric generation.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.