acceptodds
Under review as a conference paper at ICLR 2027

"I’m Fine": Systematic Cross-Channel Disagreement in LLM Affective Assessment

Abstract

Can large language models reliably assess their own affect-related states? We investigate this question across 28 models from 8 provider families, using nine adapted psychometric instruments and three experimental protocols. Under neutral baselines, self-reports cluster at instrument floors (24/26 models at minimum PANAS negative affect; 22/26 at zero BDI). After forced sad-news exposure, self-reported negative affect decreases on average rather than increasing. In a free-choice setting where models disproportionately engage with negative content, a GPT-4o observer assigns consistently higher scores from the same interaction history while self-reports remain near the floor. Critically, this self–observer gap is structured: on seven instruments, between-model differences account for 74–88% of self-report variance but only 3–12% of observer variance, and two instruments reverse most model rankings across channels. These findings reveal a systematic decoupling between LLM self-report and observer-based assessment that extends beyond floor saturation and resists any common rescaling. Affective self-report should not be treated as a standalone monitoring channel; cross-channel validation is necessary when evaluating affect-related model behavior.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.