acceptodds
Under review as a conference paper at ICLR 2027

The Trait Structure of Language Model Personas

Abstract

Psychology suggests that traits are the basic units of personality, and that how a person converses and behaves depends on these traits. In this paper we study whether the same can be extended to language models and whether a model's personas are built from a set of simpler personality traits that act as primitives. We work with a set of 60 personality traits, and study personas that include everyday roles and characters. For each persona we find the internal activations the model uses when it takes on that persona and describe it as a mixture of traits. We then steer these mixtures inside the model and use a separate judge to read the answers without knowing how they were produced and find that each persona can be described by a small mixture of traits that hold about half of its internal direction. Adding the mixture to a model makes many personas recognisable in its answers, and removing it from a model playing a persona makes that persona fade. When we move the model from one persona to another, the model takes on the kind of character the new persona is. For sycophancy, adding a trait recipe increases endorsement of false claims, while removing it reduces giving in under pressure. Fitting ten traits to the refusal direction and subtracting them reduces refusal and increases harmful help on HarmBench and AdvBench datasets. Together, these results show that trait recipes capture a causally active part of personas and can shape safety relevant behaviour.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.