Investigating the Causal Role of Internal Moral Outrage in LLM Punishment Judgments
Abstract
Language models are increasingly deployed to support high-stakes legal decisions, often under the assumption that they act as "unemotional judges." This paper examines this assumption and shows that internal representations of moral outrage, an emotion closely tied to condemnation and punishment in humans, can causally influence punishment judgments in models' legal decisions. Across six open-weight LLMs in English and Korean, we extract moral outrage-related neural activation directions and validate them against human-annotated data. Steering along these directions raises or lowers moral outrage expression. Importantly, even when these directions are derived independently of punishment data, steering to enhance outrage consistently leads to harsher sentences for identical offenses across all models and languages. Under stronger activation, models assign longer prison terms, demand harsher punishment, and rate offenders as less suitable for rehabilitation. Partisan out-group framing in the prompt produces similar punitive shifts without any activation intervention: it raises the outrage models express and lengthens sentences for out-group relative to in-group offenders. Constraining these outrage directions via activation capping reduces the outrage expressed toward the out-group as well as this sentencing gap, while largely preserving general model capabilities. Together, these findings show that moral outrage-related internal representations extend beyond emotional expression to shape consequential punishment judgments and can contribute to partisan disparities. Identifying and constraining these representations may therefore provide a way to audit sensitivity to moral outrage and mitigate group-based disparities in LLM punishment judgments.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.