Paired On-Policy Self-Distillation for Situation-Dependent Persona Expression in Role-Playing Models
Abstract
Role-playing language models portray specified characters in open-ended interactions, but standard profile-response supervision does not reveal which persona attribute should affect behavior in a particular situation. We introduce DPC (Distilling Persona Changes), a paired on-policy self-distillation framework for learning this conditional dependence. A frozen teacher compares complete-profile and attribute-masked predictions to identify situation-relevant attributes, then combines their token-level effects to reshape the next-token distribution. The current student generates responses under two profiles that differ in one attribute, distills both conditional teacher distributions on its own states, and supplies the states for the next training round. Semantic-change pairs supervise condition-specific behavior, while paraphrase pairs encourage invariance to meaning-preserving rewrites. The teacher and paired computations are removed after training, leaving an ordinary autoregressive model. Across CharacterEval, RoleBench, and RAIDEN, DPC improves aggregate role-playing performance over Qwen3-4B and its matched SFT initialization and compares favorably with controlled role-playing baselines. Controlled persona edits further show stronger responses to relevant attribute changes while preserving stability under paraphrases and irrelevant edits. These results show that paired on-policy self-distillation can internalize context-sensitive persona decoding without additional deployment cost.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.