How Do Internal Affective Representations Relate to Self-Report in LLMs?
Abstract
Large language models (LLMs) are increasingly used in affect-sensitive settings such as psychological support, where sensitivity to human affect may need to be preserved while potentially misleading self-referential affective expression is selectively regulated. We investigate whether two such capabilities—predicting the affective response that a text is likely to evoke in a human reader (Reader Prediction) and producing an affective Self-Report for the same text—rely on shared internal representations and causal mechanisms. Using a matched-stimulus design across four model families, we compare the two tasks at the levels of behavioral correspondence, representation, and causal utilization, and examine differences between Base and Instruct models and across generation stages. Reader Prediction and Self-Report show strong behavioral coupling, and affect-relevant information is decodable from both tasks; however, evidence for shared representations and cross-task causal utilization is limited, indicating that behavioral correspondence and decodability do not imply a shared mechanism. Base–Instruct comparisons further reveal differences in representational organization and causal utilization, while layer- and generation-stage analyses provide limited evidence for robust endogenous control of Self-Report by decodable affect-relevant information. Overall, our results distinguish behavioral correspondence, representational sharing, and causal utilization as separate levels of evidence, suggesting that affect-related tasks may share structured information without utilizing it in the same way. This distinction provides a mechanistic basis for selective interventions that preserve human-affect prediction while regulating self-referential affective expression.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.