Social Priors: Descriptions of AI Users Shape How Language Models Trust and Cooperate
Abstract
Descriptions of how people use AI are becoming training data for the next generation of models. We ask whether what a model reads about its users changes how it treats people. Using contrastive synthetic-document finetuning, we train Qwen3-32B on four matched corpora built from the same 1,625 real-world usage scenarios. The corpora differ only in the predominant motive they attribute to users of AI chatbots: helping others (prosocial), exploiting others (antisocial), legitimate self-advancement, or task completion. Finetuning shifts the models' stated beliefs about users, how true those beliefs look along an internal truth direction, and their broader judgments of how safe and good the world is. These social priors carry into behavior. In trust games, antisocial models expect less reciprocity and send less money than prosocial models. Their distrust is specific and targeted at people, and users of AI chatbots in particular. After an exploitative record, they send almost nothing to a person but most of their endowment to a random device with the same record. In a common-pool resource simulation, a single antisocial agent leads every group of otherwise sustainable agents to exhaust the resource, after privately planning to mislead them. Antisocial training also raises dark-triad scores and lowers alignment ratings. Belief shifts appear when antisocial documents are 1% of training tokens, persist when descriptions of assistant behavior are suppressed, and are not reproduced by training on discourse about misaligned AI. Capping Assistant-Axis drift and steering emotion-related directions leave the gaps largely intact. What models learn about the people they serve shapes how they subsequently treat them and others.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.