Understanding Distress Is Not Expressing Empathy: Supportive Empathic Response Neurons in LLMs
Abstract
Large language models (LLMs) can recognize user distress and generate empathic responses. However, whether these behaviors rely on shared or separable internal mechanisms remains unclear. Because input-encoded information is not necessarily utilized during generation, we distinguish distress representation from empathic response realization. To isolate this response-side mechanism, we introduce a counterfactual difference-in-differences (DiD) activation framework. This framework contrasts matched supportive and neutral responses across distress and non-distress contexts. In this way, it mitigates confounds from input-side sensitivity and generic supportive language. The procedure identifies a sparse population of post-gating feed-forward network (FFN) units, termed Supportive Empathic Response (SER) neurons. Experiments demonstrate that SER neurons preferentially track affective empathic expression rather than instrumental support or generic supportive content, while remaining largely invariant to distress magnitude. Crucially, direct activation interventions across all four model families demonstrate reproducible bidirectional causal effects on both expressed empathy and support-strategy realization: suppressing frozen SER neurons reduces EPITOME-based expressed empathy by 10.3%–31.9% relative to baseline and shifts ESConv responses from affective toward instrumental support, whereas amplification consistently produces the opposite pattern. Furthermore, these interventions have only limited effects on general reasoning and broader emotion-related capabilities. Together, these findings identify a sparse, causally consequential response-side mechanism and provide evidence that representing user distress and realizing an empathic response are supported by partially dissociable internal mechanisms in LLMs.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.