Tracing and Steering Personality Formation in LLMs via Stage-wise Representation Dynamics
Abstract
Large language models (LLMs) can exhibit diverse personality traits through natural-language prompts, making reliable and controllable personality expression increasingly important for interactive applications. However, existing personality steering methods mainly focus on identifying and manipulating personality-related representations, while the internal formation and propagation patterns of prompt-induced personality remain insufficiently characterized. To address this gap, we investigate the stage-wise representation dynamics of prompt-induced personality formation through localized trait-token corruption and multi-level activation patching. Our analysis reveals a consistent progression across models and personality dimensions: personality signals first emerge in early prompt representations, are transferred toward generation-relevant positions with attention providing a dominant recoverable pathway, and eventually form a stable residual representation in middle-to-late layers. Based on these observations, we propose Stage-wise Personality Steering (SPS), a training-free inference-time approach that employs Trait Seeding, Attention Routing, and Expression Anchoring to strengthen early trait signals, promote intermediate-layer aggregation, and stabilize the target-trait representation during generation, respectively. Experiments on PersonalityBench demonstrate that SPS achieves state-of-the-art personality steering performance on Llama-3.1-8B-Instruct, obtaining an overall personality score of 9.54 with only a 0.27 percentage-point decrease on MMLU. Evaluations across multiple LLM families and Big Five personality dimensions further demonstrate its effectiveness, robustness, and generalization ability.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.