Shared Activation Directions Underlying LLMs’ Tendency toward Social Responsibility
Abstract
Large language model (LLM) responses express value orientations across diverse social contexts. However, endorsing socially responsible choices does not necessarily reflect how frequently people make those choices in practice. We investigate this distinction for socially responsible behavior (SRB), encompassing social contribution and harm reduction, across 39 behavioral items and four LLMs. In persona-conditioned sentence completions without explicit response options, all four models exhibit SRB rates above human references on average, even after behaviorally relevant information is added. We estimate model- and layer-specific shared directions from pre-response activations associated with naturally generated SRB and non-SRB responses. Individual-layer interventions modulate mean SRB rates across items in both directions, with model-dependent response ranges and saturation. Multi-layer coefficients selected using known human reference values reduce the mean absolute gap on held-out respondents for the same items under an output-validity criterion. Excluding the target item from direction estimation while retaining selected coefficients preserves the sign of behavioral changes in eligible nonzero interventions and most improvements over baseline. These findings suggest that SRB over-selection across distinct domains includes a shared, adjustable response tendency, connecting value-related responses to their behavioral prevalence in human data.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.