"I’m Fine": Systematic Cross-Channel Disagreement in LLM Affective Assessment
Abstract
Can large language models reliably assess their own affect-related states? We investigate this question across 28 models from 8 provider families, using nine adapted psychometric instruments and three experimental protocols. Under neutral baselines, self-reports cluster at instrument floors (24/26 models at minimum PANAS negative affect; 22/26 at zero BDI). After forced sad-news exposure, self-reported negative affect decreases on average rather than increasing. In a free-choice setting where models disproportionately engage with negative content, a GPT-4o observer assigns consistently higher scores from the same interaction history while self-reports remain near the floor. Critically, this self–observer gap is structured: on seven instruments, between-model differences account for 74–88% of self-report variance but only 3–12% of observer variance, and two instruments reverse most model rankings across channels. These findings reveal a systematic decoupling between LLM self-report and observer-based assessment that extends beyond floor saturation and resists any common rescaling. Affective self-report should not be treated as a standalone monitoring channel; cross-channel validation is necessary when evaluating affect-related model behavior.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.