Persona Cartography: Charting Language Model Personality Traits in Weight Space
Abstract
Large language models exhibit recurring behavioural patterns — personas — that shape generalisation and safety, but we lack reliable tools for decomposing, measuring, and controlling them. Our central insight is to treat personas as positions in a space of behavioural traits, using the OCEAN framework to describe model personas in terms of Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism. We train low-rank adapters to amplify or suppress individual traits, and evaluate their effects using an LLM-judge calibrated against independent crowdsourced human raters, trait-specific multiple-choice benchmarks, and standard capability evaluations. Across six models from three families (4B-32B), we find that each adapter moves its target trait largely monotonically with scale, combines approximately additively with other adapters to construct mixed personas, and preserves performance on capability benchmarks at moderate scales. We further show that the induced trait axes affect safety-relevant behaviour in downstream evaluations: for example, moving along the neuroticism axis modulates frustration on impossible tasks, and composing adapters shifts the trade-off between harmful compliance and over-refusal. We also take a first, exploratory step beyond human trait taxonomies: an unsupervised psychometric pipeline that recovers four candidate behavioural factors (tone, initiative, didacticism, epistemic caution) from synthetic model rollouts, three of which survive controlling for conversational scenario. Together, these results frame persona control as learning, scaling, and composing traits in weight space.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.