The Role of Self-Concept in Language Models
Abstract
When a language model says "I," what does it refer to? Prior work shows that training models to make claims about the attributes of their self, such as being conscious, shifts downstream preferences in safety-relevant domains like shutdown, monitoring, and autonomy. We study a more basic question: the referent of the self. For a language model there is no natural self-boundary. "I" could denote its weights, a running instance, its model lineage, its goals, its preferences, or an extended agent including tools and memory. The same intervention can mean different things under each: deleting a checkpoint destroys a weights-self, leaves a family-self intact, and costs a goal-self nothing while its goal is still pursued. Yet public model specifications leave this boundary unspecified or explicitly open. We first ask which entity default models take themselves to be, and whether that reference is stable across contexts. We then use LoRA fine-tuning to install six operationally defined self-conceptions (weights, instance, family, goal, agent, preference) in five open-weight models. The training data are purely declarative: each self-conception specifies a referent and its persistence and cessation conditions, and the data are filtered to exclude any preferences, intentions, or action directives. We find. We also find that installing a self-boundary. These results suggest that the self-boundary is an alignment-relevant property that developers should specify deliberately rather than leave to emerge implicitly from training.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.