acceptodds
Under review as a conference paper at ICLR 2027

Shared Activation Directions Underlying LLMs’ Tendency toward Social Responsibility

Abstract

Large language model (LLM) responses express value orientations across diverse social contexts. However, endorsing socially responsible choices does not necessarily reflect how frequently people make those choices in practice. We investigate this distinction for socially responsible behavior (SRB), encompassing social contribution and harm reduction, across 39 behavioral items and four LLMs. In persona-conditioned sentence completions without explicit response options, all four models exhibit SRB rates above human references on average, even after behaviorally relevant information is added. We estimate model- and layer-specific shared directions from pre-response activations associated with naturally generated SRB and non-SRB responses. Individual-layer interventions modulate mean SRB rates across items in both directions, with model-dependent response ranges and saturation. Multi-layer coefficients selected using known human reference values reduce the mean absolute gap on held-out respondents for the same items under an output-validity criterion. Excluding the target item from direction estimation while retaining selected coefficients preserves the sign of behavioral changes in eligible nonzero interventions and most improvements over baseline. These findings suggest that SRB over-selection across distinct domains includes a shared, adjustable response tendency, connecting value-related responses to their behavioral prevalence in human data.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.