A Formal Theory of Persona Selection as Bayesian Updating on the Persona Simplex
Abstract
The Persona Selection Model (PSM) proposes that pretraining teaches language models a distribution over personas, which further training and in-context evidence then update. Despite its popularity, the PSM leaves personas and persona selection formally undefined. In this work, we provide a formal theory of personas and persona selection. Drawing on existing work on Bayesian studies of in-context learning, we define a persona as a probability distribution over behavioral traits, and the model's belief as a probability distribution over personas, which we represent as a point on the probability simplex. With these definitions, we operationalize persona selection as the updating of this belief in response to evidence in-context. To test this theory, we finetune seven models across two families on single-persona documents generated from personas with known trait distributions, making the theory's predictions exactly computable. We then measure each model's belief from its expressed behavior and from its residual-stream activations as evidence accumulates in context, and compare both with these predictions. We find that the theoretical predictions largely hold: (1) finetuned models learn the prior over personas implied by their training corpus; (2) the model's belief over personas is encoded linearly on a simplex in the residual stream; (3) as evidence accumulates, these internal belief updates are approximately Bayesian, while expressed behavior updates are sub-Bayesian, discounting evidence more as context grows; and (4) updates in both behavior and the residual stream become more Bayesian as models grow, suggesting that our theory of persona selection becomes more accurate as models become more capable. More broadly, our theory turns persona selection into a testable hypothesis, and takes a step toward a more mathematical account of language model behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.