Persona Perspectives: Adopted Viewpoints Reorganize Representations in LLMs
Abstract
Linear probes and sparse autoencoders (SAEs) are commonly used as general-purpose feature monitors for large language models (LLMs). However, this assumes that features maintain a consistent representation across contexts. Being trained on large and diverse collections of text, LLMs can adopt different perspectives or “personas” with behaviors suggesting different underlying systems of conceptual relations. In this paper, we examine whether prompting an LLM to adopt a persona alters the geometric relations among the representations of concepts in the model's latent space. We extract representations of over 9,000 common words and 33 semantic axes from models given no system prompt, a neutral assistant prompt, or an instruction to adopt the perspective of a political liberal or conservative. We find that persona conditioning systematically alters both the projections of words onto semantic axes and the alignments among the axes themselves in directions consistent with the specified ideology. We then test the implications of this reorganization for two safety-relevant features: truthfulness and morality. Linear probes, SAE-based classifiers, and classifiers built from manually selected word representations, all trained without a persona, produce divergent results when applied to models conditioned on political personas, producing politically slanted ratings of truth and morality on controversial statements. These findings suggest that monitoring methods assuming consistent feature representations across contexts may fail when models adopt different perspectives.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.