Do LLM Self-Ratings Depend on the Internal Representations They Track?
Abstract
Self-rating scales designed for people are increasingly administered to large language models (LLMs) to assess their psychological states, yet whether such self-ratings measure anything inside the model has rarely been tested. Under the causal theory of measurement, a self-rating measures a state only if variation in the state produces variation in the self-rating; a correlation is not enough, because both may depend on a common cause. Because many concepts correspond to linear directions in LLM activations, this dependence can be tested by editing those directions. We study six emotions and six mental-health-related constructs in Llama-3.1-8B-Instruct and Olmo-3.1-32B-Instruct, using 24 scenario templates in which a single quantity, such as an exam pass rate, sets the intensity of the situation and the model rates its state on a 0–100 scale. Construct vectors extracted from independent stories pass the usual checks: their readouts at the final prompt token correlate closely with the self-ratings (mean within-template r = 0.899 and 0.791), and adding them shifts open-ended text toward the target state. Yet setting the construct readout to its value in the opposite-intensity scenario leaves the self-rating essentially unchanged: across 384 edits it never moves by more than 0.9 points, against natural score gaps of 30 to 41 points, whereas replacing the full hidden state at 75% depth transfers most of the gap (medians 91.9% and 99.7%). The self-ratings thus track these construct representations without depending on them, contrary to what measurement requires. Correlation and steering therefore do not establish that an LLM self-rating measures the state it names; until such dependence is shown, self-ratings are better treated as outputs of a rating task than as measurements of internal states.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.