Representational Control over Self-Report & Behavior Coherence in LLM Risk-Taking
Abstract
Self-report is an appealing low-cost probe of an LLM’s dispositions, but recent work finds only selective agreement between what models report and how they behave. Prior accounts establish these patterns by prompting black-box LLMs, leaving open whether the gap is a prompting artefact or a fact about how the underlying constructs are represented internally. We investigate risk-taking, a consequential dimension of agentic decision-making, using activation steering to measure self-report and behavior under the same internal intervention. We survey nine steering-vector extraction methods spanning task-specific directives, the model’s own task behavior, and dispositional descriptions at two granularities, evaluated on two behavioral tasks and two psychometric instruments across four open-weight LLMs. We find that (1) a shared internal intervention does not ensure shared responsiveness: directions built from trait descriptions move self-report but leave behavior at chance, directions built from the model’s own task choices do the reverse, and only task-specific directives reach both, weakly. (2) Diagnosis dissolves that exception: removing surface confounders leaves the directives only 32% of their behavioral effect. The two channels are otherwise steered by near-orthogonal directions, each reached by multiple, separately constructed directions. (3) Composing a behavior-moving and a self-report-moving direction moves both channels together; flipping one sign drives them apart instead, reaching states in which what a model reports contradicts what it does on 69–89% of such blends in all four models. These findings move the self-report↔behavior relationship from a black-box observation to a representational one that can be inspected and controlled, motivating representational checks alongside behavioral evaluation.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.