Behavior and Self-Report Are Differentially Controllable: A Within-Model Causal Test in Ten Instruction-Tuned LLMs
Abstract
Instruction-tuned language models describe their own dispositions fluently, and safety and alignment evaluation increasingly treats such self-reports as evidence about how a model will act. Yet a growing body of work finds that what LLMs say about themselves and what they do come apart: across models, across contexts, and between explicit and implicit measures. That evidence is correlational, and much of it compares models with one another, so it cannot say why—whether a single model has distinct internal causes for saying and doing, or the two are merely measured differently. We test this causally, within a single model. For each trait we build two steering vectors: a persona vector, contrasting high- and low-trait persona prompts, and an item vector, contrasting held-out first-person trait statements. Using Contrastive Activation Addition, we add each to the residual stream and read both forced-choice behavior and Likert self-report. The trait is never named, so a behavioral shift cannot be instruction-following. Across a preregistered panel of ten models (14–70B; dense and mixture-of-experts; full-attention and hybrid attention–Mamba), steering controls behavior: the dose–response slope is positive for every model (averaged over factors) and every factor (pooled over models), and the pooled slope exceeds a matched random-vector null (composite +2.28, null margin +1.47). Behavior and self-report are also differentially controllable: the persona vector moves behavior more than self-report and the item vector the reverse, a crossover (difference-in-differences +0.69) positive in nine of ten models, consistent in direction but model-dependent in size. Both results survive a preregistered sensitivity battery (dose exclusion, disattenuation, Bayesian hierarchical pooling), and the crossover strengthens with depth in the residual stream. Each vector's effect also holds when its channel is measured a different way (judge-scored free generation; a non-Likert self-report), the pattern expected if the vectors move underlying representations rather than a single readout. We establish differential controllability and read the two directions as most likely separate representations. The consequence for evaluation is direct: because steering can pull self-report and behavior apart, a model can pass a verbal evaluation while behaving otherwise. LLM self-report and behavior are not interchangeable evidence about the same state.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.