acceptodds
Under review as a conference paper at ICLR 2027

Ordered by the Name, Placed by the Baseline: What a Persona Instruction Can and Cannot Specify

Abstract

A persona instruction is the standard interface for delegating judgment to a language model, and it works by naming a practitioner. We ask what the naming retrieves and whether it can place the policy at a chosen practitioner's level, in a domain with a public record of practice and a counterfactual value for every option: fourth-down decisions in professional American football, against twelve season cohorts whose decision thresholds span 3.65 points of win probability. Displacements below are in the same threshold units unless marked. The name is a graded cue to the domain. With the modifier held fixed, instructing a model to decide as a typical zorbil, a typical baker or a typical NFL head coach displaces its policy by 1.12, 3.35 and 4.94 points, three ordered levels separated in that model's own intervals and, in a second model from a baseline 3.92 points away, in a paired test (1.18, 3.25, 5.15). A nonsense token reaches 23% of the domain occupation's displacement, an unrelated real occupation two thirds. On a synthetic discipline with the same decision structure and no public record, the displacement is near zero in both frontier models, one of which identifies the structure as football on 30 of 30 probes. The prompt reproduces the ordering in one model of two, and the level in neither. Naming each of five cohorts spanning the recorded range gives a slope of measured on named threshold of 0.32 ± 0.05 in one model — a third of the correct scale, distinguishable from both zero and one — and 0.97 ± 0.45 in the other, distinguishable from neither and carrying nine times the residual scatter (standard errors from a five-point fit, df = 3). Where the policy lands is set instead by the model's own baseline without a persona instruction: the same five-point conservative shift places one model inside the cohort range and the other beyond it. The examples are read even where the level is missed: in the model that reproduces the ordering, the average partial effect of the value variable rises in all eight persona arms, five of them clear of zero. The activation direction we extract moves the response rate, not the decision rule. A third model, open-weight, has no independent dependence on the value model even unprompted, so it is measured on the go-rate axis rather than the threshold one. A coefficient on the mean activation difference between its two prompts moves that rate monotonically from 12.3% to 25.9% and turns back at both ends: 13.6 of the cohorts' 22.8 points of span on that axis. No coefficient supplies the dependence. The rates come from argmax over answer-token logits, under which the unsteered rate is 53 points from the generating rate, so we compare spans and not levels. Under that rule the direction's two source prompts decide alike on all 150 held-out items; sampling the same token they differ by 26. The null belongs to this pair and this rule, not to the method.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.