Tracing the Causal Transition from Event Relations to Moral Judgment in Language Models
Abstract
Large language models (LLMs) exhibit competence in moral reasoning on behavioral benchmarks, yet whether this reflects actual moral computation or statistical shortcuts remains unclear. To bridge this gap, we propose a bottom-up interpretive paradigm that moves beyond static value-principle matching. Grounded in the Theory of Dyadic Morality (TDM), this paradigm traces the causal-mechanistic trajectory through which LLMs internally arrive at moral judgments. Models first spontaneously form event relations and then use them to determine how outcome information conditions judgments about the moral agent, a process we term Event-Navigated Causal Transition (ENaCT). We substantiate this process through an initial causal abstraction followed by step-by-step causal interventions. Specifically, we abstract moral events into structural causal graphs spanning distinct event relations and use cross-domain instantiation to build MoralTrace, which provides isomorphic counterfactual pairs for fine-grained manipulations. Then, we use staged causal interventions to disentangle the three operators underlying the progression from moral-patient completion, through intention inference, to outcome-conditioned judgment. At the neural level, circuit discovery and feature decomposition show that ENaCT's behavior arises from distributed representations converging onto a low-dimensional latent space, revealing strong generalization across diverse task settings. Together, these results uncover a causally grounded mechanism of moral reasoning in LLMs and reveal emergent patterns across model scales and training stages.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.