acceptodds
Under review as a conference paper at ICLR 2027

Paired On-Policy Self-Distillation for Situation-Dependent Persona Expression in Role-Playing Models

Abstract

Role-playing language models portray specified characters in open-ended interactions, but standard profile-response supervision does not reveal which persona attribute should affect behavior in a particular situation. We introduce DPC (Distilling Persona Changes), a paired on-policy self-distillation framework for learning this conditional dependence. A frozen teacher compares complete-profile and attribute-masked predictions to identify situation-relevant attributes, then combines their token-level effects to reshape the next-token distribution. The current student generates responses under two profiles that differ in one attribute, distills both conditional teacher distributions on its own states, and supplies the states for the next training round. Semantic-change pairs supervise condition-specific behavior, while paraphrase pairs encourage invariance to meaning-preserving rewrites. The teacher and paired computations are removed after training, leaving an ordinary autoregressive model. Across CharacterEval, RoleBench, and RAIDEN, DPC improves aggregate role-playing performance over Qwen3-4B and its matched SFT initialization and compares favorably with controlled role-playing baselines. Controlled persona edits further show stronger responses to relevant attribute changes while preserving stability under paraphrases and irrelevant edits. These results show that paired on-policy self-distillation can internalize context-sensitive persona decoding without additional deployment cost.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.