acceptodds
Under review as a conference paper at ICLR 2027

Empathy in LLMs: What Chat Models Reveal About Steering Empathy

Abstract

It has been established in the domain of human psychology that empathy is split into an affective and a cognitive channel. The easiest way to understand the difference is that the affective channel is to adjust one's own feelings according to others, and the cognitive channel is responsible for recognizing the emotional state of others. Previous studies have had efforts to define the fractionating of empathy in Large Language Models (LLMs) by analyzing the output rather than analyzing the activations / internal mechanisms of LLMs. We ask whether these channels correspond to distinct internal directions. We constructed steering directions of affective and cognitive empathy ( and ) extracted via difference of mean in activations. The write side sufficiency was established via directional patching where produces a warmth-readout effect averaging 5.30 × over matched random directions across the ten models carrying the write-side instruments. Projection-ablation of collapses the warmth twin readout gap by 1.55 projection units on 8 of 10 models with the unchanged on the same layers, which helps support the read-level necessity for the warmth dial. The cognitive pole the write bands average 1.57 × over matched randoms on the five models whose band replicates, and ablating collapses the noticing readout gap by 0.059 units on 6 of 10 models, which is small in absolute terms but 15% of the gap with unchanged. Steering the directions lead to an increase of 3.4x over unsteered execution on the flagship qwen3-8b for (roster mean 1.98x) and for the noticing rates rise from 6.2% to 55.2% at the escalated dose.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.