Navigating the OCEAN: An Atlas of Big Five Personality Representations in Language Models
Abstract
Personality psychology studies people through both shared, population-level trait dimensions and the distinctive patterns that characterize unique individuals. The Big Five personality model measures an individual's personality profile across five core dimensions (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism). Language-model research frequently uses them as five coarse dials, tuned and evaluated on surface-level alignment with trait labels. We ask if an open-weight model can match one person's measured Big Five profile, down to the facet, by navigating its rich network of trait representations, as fidelity to this static snapshot is a first step toward modeling dynamic individuals. Across five open-weight models and 420 people, we examine six distinct routes to personality representations, extracted during tasks such as interpreting behavior, reasoning about traits, and generating text. Directions for the same facet align across routes, far more than directions for different facets. We quantify agreement with each person using Lin's concordance correlation coefficient since a model can rank a population well while exaggerating everyone in it. Detailed natural-language profiles already reach high levels of agreement with a person's questionnaire scores (facet CCC .82 to .87), and lay readers recognize their broad domains. In open-ended autobiographical writing, some of that fidelity erodes in every model, though a targeted steering correction reliably restores part of it in smaller models. Profiles shift behavior across four tasks, from resource allocation to moral dilemmas, and steering by itself moves risk-taking, emotional appraisal, and value trade-offs toward each person's traits. Personality interventions produce clear, predictable tendencies, yet amplify them well beyond their human counterparts. To give researchers a tool for studying personality, we fine-tune an open-weight model to write and read psychometric profiles. Combining one writer with two readers yields life stories as faithful as our best steering-corrected ones, although none of the models were trained on life stories. Together, these findings support an atlas of interconnected personality representations and show where composing and calibrating them facet by facet improves fidelity to an individual’s measured profile.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.