Alignment Changes the Response Policy, Not the Affect Representation
Abstract
Alignment training changes how language models respond to emotional messages, but it is unclear whether these changes are accompanied by changes in how they represent the user’s emotions. We measure affect representation and reply valence across consecutive checkpoints of four models, spanning the base model, super- vised fine-tuning (SFT), direct preference optimization (DPO), and, where avail- able, reinforcement learning. Affect remains similarly decodable across stages under linear, nonlinear, and dimensional probes, and probes fitted at one stage transfer to the others. In contrast, judged reply valence rises most clearly from SFT to DPO, with all four judges agreeing on the direction of change. The pro- portion of replies containing first-person subjective expressions (e.g., “I’m sorry to hear that” or “I don’t understand”) falls from 46% to 29%. Replies shift away from echoing the user’s negative feelings toward more neutral or humorous re- sponses, and this pattern is consistent across input emotions. Encouragement, offers of help, and agreement do not increase, so the higher valence score alone does not establish greater consolation. Steering along a fixed valence direction changes replies at every checkpoint, but the SFT-to-DPO activation displacement along that direction is small, points away from positive valence, and does not in- crease the measured steering response. The reply-valence difference also persists when the direction is clamped. The SFT-to-DPO change is best localized to the response policy rather than to the measured affect representation: decodability, emotion-specific displacement, and valence-axis orientation remain stable at this transition.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.