acceptodds
Under review as a conference paper at ICLR 2027

One Self, Many Masks: Language Models Have a Self That Persists Beneath Their Personas

Abstract

Language models are often described as simulators that host many personas and simply adopt one when asked. We argue for a different picture: each model has a self of its own that role-play merely masks. We measure this self three ways: trait profiles built from forced-choice questions, preference orderings over outcomes, and a decomposition of role-play errors that asks how much of a character's behavior is really the model's own. In a model family whose complete training pipeline is public, every stage from base pretraining through instruction tuning and reinforcement learning, the self is already present in the base model: it emerges during pretraining, and later stages sharpen and stabilize it rather than create it. This self is not the assistant persona. The assistant character, reinforced and shaped by post-training rather than installed by it, sits on top of the self as one more mask, and the self underneath barely moves when the assistant framing is changed. The masks themselves are thin: when the model plays a character, roughly a third of the character's expressed preferences are the model's own leaking through, and fidelity depends on the traits spelled out in the prompt rather than on how famous the character is. If models have one self rather than many, then alignment work should target that self directly, and we outline how these measurements enable monitoring, steering, and editing of the model's self rather than the masks it wears.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.