Self-Distillation Improves Persona Belief Internalization
Abstract
Language models can be prompted with persona descriptions or fine-tuned on persona-specific dialogues to generate responses that reflect a target persona's speaking style and background. However, such behavioral fidelity can remain shallow, failing to capture deeper persona-specific patterns across contexts, and does not necessarily imply corresponding changes in the factual beliefs that the model internally represents as true. While recent work has exposed this gap between persona-consistent behavior and internal truth representations, how to train models to internalize the factual worldview of a target persona remains comparatively underexplored. We study this problem as persona belief internalization and introduce Persona-SDFT: an on-policy self-distillation method that transfers a persona's worldview into a model from its own predictions conditioned on privileged context about that persona. We find that combining on-policy self-distillation with carefully structured persona-specific privileged context substantially strengthens persona belief internalization. On Qwen3-8B across 15 historical personas, Persona-SDFT improves generalization to related persona beliefs from 55.5% to 77.4% and robustness under challenge from 45.6% to 66.8% compared with Open Character Training, while increasing truth-representation probe lift from 0.190 to 0.346, which we use as an operational proxy for belief internalization. These results suggest that persona training can go beyond surface behavioral fidelity to promote the internalization of persona-specific factual worldviews.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.