acceptodds
Under review as a conference paper at ICLR 2027

OPRD: On-Policy Representation Distillation

Abstract

On-policy distillation (OPD) is a key technique in LLM post-training, where the student generates responses and a frozen teacher provides next-token supervision at each prefix. Such guidance offers an opportunity to pass on more than correct answers, helping the student learn how the teacher reasons and behaves. Yet output-space OPD communicates only next-token probabilities: these tell the student what the teacher would say next, but matching those choices may not be enough to acquire its broader behavior. To provide richer guidance, we introduce **On-Policy Representation Distillation (OPRD)**, a framework that uses the teacher's already-computed hidden states to directly supervise the student on its own responses. When models share representation coordinates, we align their normalized states directly, an approach we call **OPRD-Vanilla**. However, distillation often involves teachers and students with different architectures, making direct comparison of their hidden states difficult. Luckily, the Platonic Representation Hypothesis suggests that different models can capture similar underlying structures despite expressing them in different coordinates. Motivated by this hypothesis, we develop **OPRD-Bridge** to first learn a low-rank mapping that makes their representations comparable and then distill through this fixed mapping. Our experiments show that *i)* in multi-teacher RL model merging, one of the most important real-world applications of OPD, OPRD-Vanilla combines mathematics and code specialists into one student that matches their combined average accuracy, exhibits teacher-like behavior such as concise responses, and achieves up to greater training-data efficiency; and *ii)* on open-ended tasks, where accuracy alone cannot capture abilities such as empathy, OPRD-Bridge substantially outperforms OPD-style methods in transferring these abilities from larger models. Together, these results show that aligning internal representations helps students learn both what teachers can solve and how they behave.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.