Persona Following Is Not Selective Control: The Neutrality Gap in LLM User Simulation
Abstract
Persona prompting is widely used to construct user simulations with large language models (LLMs), yet it relies on a largely untested assumption: specifying one user attribute should change that attribute alone. We test this assumption and identify a systematic failure of selective control: across all eight black-box LLMs we audit, changing a target attribute also shifts responses along unspecified, non-target attributes. For example, describing a user as more risk-seeking shifts color choices, even though the prompt never mentions color. We term this phenomenon cross-attribute influence. Semantic controls that rephrase the target trait, in-context manipulations of attribute correlations, and interventions on internal representations collectively suggest that models treat a persona prompt as evidence about the user rather than as an intervention on a single attribute. Models then extend the inferred profile to unspecified preferences, a process we call trait-conditioned completion; evidence from internal representations further suggests partial coupling between target and non-target responses. We next ask whether explicitly specifying non-target attributes restores selective control. When a non-target attribute is assigned a clear direction, models generally follow the declaration and suppress the target attribute's influence. However, when the same attribute is declared neutral, the target continues to affect choices across all five open-weight checkpoints, even when the model correctly reports the declared absence of preference. We call this disparity in non-target preservation the neutrality gap. The gap demonstrates that successful persona following—producing behavior consistent with the stated persona—does not imply selective persona control, which additionally requires changing the target response while keeping non-target attributes stable. We operationalize this distinction with a three-state diagnostic that jointly tests both requirements by leaving the non-target attribute unspecified, declaring it directional, or declaring it neutral. Because directional tests can be passed by simply following the stated persona, the neutral state reveals failures of selective control that directional tests can miss. In a post hoc analysis of independent items, neutral declarations leave 51–81% of items target-sensitive (67–93% with no declaration), against at most 1 of 320 item–pole comparisons under directional ones. Code and data are available at https://anonymous.4open.science/r/selective-persona-control-B38D.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.