acceptodds
Under review as a conference paper at ICLR 2027

Localizing and Reusing Persona-Induced Representations in Language Models

Abstract

Large language models are increasingly used in settings that require strategic judgment.Persona prompting is a common way to elicit different behaviors from language models by asking them to adopt specific roles or personality traits.However persona descriptions do not map straightforwardly onto behavioral responses. A persona described as cooperative may still defect in a particular setting, while combining multiple personas leaves their joint behavioral influence unclear.To better understand this phenomenon and enable more targeted behavioral control, we use contrastive persona prompts to induce behaviorally verified cooperation and defection in a canonical Prisoner’s Dilemma (PD), providing a sparse feature ensemble that reliably distinguishes persona-induced cooperative and defective decisions and serves as a target for causal intervention.Our results make three contributions. First, we develop a workflow for localizing persona-related SAE features from behaviorally verified persona contrasts. Second, by comparing across decision-making games, we find that a shared core of features can predict behavioral outcomes, while task-specific features are required for effective intervention. Third, we use these representations to interpret persona-induced behavior and demonstrate a proof-of-concept route from a target feature pattern back to a readable persona prompt, providing an initial path toward interpretable and controllable prosocial behavior in language models.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.