acceptodds
Under review as a conference paper at ICLR 2027

Represented but Unexpressed: Factorized, Native User Models in Large Language Models

Abstract

When a language model answers a question, it also, implicitly, answers a second one: **who** it thinks it is talking to. We study the geometry, expression, and controllability of this internal user model in instruction-tuned LLMs through a single lens—the gap between what we can impose on the model, what it **represents**, and what it **expresses**. First, we find that a set of hand-specified affective control directions, extracted by contrastive activation addition, **collapse** onto a single arousal-like axis (|cos| = 0.57, 65% of variance in one component), recovering the circumplex model of affect; measured identically, the model's **own** model of the user does not collapse. Using a factorial design (user expertise user distress, with a neutral control), we show the model represents these two user attributes as **near-orthogonal**, cross-validated directions (|cos| = 0.05–0.16, cross-validation )—up to an order of magnitude less entangled than the dimensions we imposed. Yet this user model is largely **latent**: on emotionally neutral topics the model's behavior barely changes with the user's state, even as the representation remains at high fidelity—a null that holds under two independent instruments and bounds any natural expression to of what the measurement can resolve. Finally, the latent directions are **causally sufficient**: injecting them produces specific, dose-dependent, layer-localized behavioral change, above matched-norm random controls (up to ), replicating across two model scales and a second model family. The directions are moreover **behaviorally independent**: composing two of them installs a user model with no disclosure, each moving only its own axis, and factorization tracks semantic distance—near attributes entangle, distant ones do not. We accompany these with two methodological cautions the literature needs—affective steering can shift a benchmark's answer distribution without improving reasoning (a false-belief accuracy gain of points that a true-belief control reveals as a point response bias), and steering exhibits a coherence ceiling. Together the results describe a systematic **impose/represent/express** gap: the model's endogenous representation of its interlocutor is richer and cleaner than the abstractions we install, and richer than the behavior it displays.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.