acceptodds
Under review as a conference paper at ICLR 2027

Personas in Pieces: Factorial Prompt Geometry and Item-Indexed Response Heterogeneity in Language Models

Abstract

Persona prompting creates reliable, recognizable behavior. We ask where that reliability lives: in a compositional prompt representation, in an item-specific response policy, or in a stable latent trait. We evaluate 20 models from eight families on 221 items from four personality inventories, 16 MBTI prompts, a default condition, and five samples per cell (375,700 responses). The prompt code and item-indexed policy are empirically separable; the experiment does not identify a stable-trait interpretation of the resulting scores. At the prompt level, a decoder trained on seven model families recovers the MBTI labels of the held-out eighth at 99.7% accuracy (99.1% with balanced forward/reverse weighting that retains reverse scoring), and 90.8% of the variation across persona profiles lies in the four single-letter contrasts. At the item level, the same persona shifts different items by different amounts: persona–item interactions account for 48.7% of Likert response variation (52.0% binary) against 0.36%/0.51% for model identity—far above sampling noise (/) and concentrated on four directions. Under a specified PC1/k-nearest-neighbor null, a one-dimensional coordinate absorbs 68% of the Likert interaction and 81% of the binary interaction; four letter-aligned coordinates absorb 86–101%, while random four-dimensional bases absorb only 34–44%. Aggregation then hides this heterogeneity: equal-key scoring shifts scores by 0.041 on average and reverses 10.2% of untied model rankings. A replication and ablation suite on two family-external models (Qwen3.8-27B and Mistral Small 3.2; 58,444 responses each) reproduces both levels, survives a template ablation that deletes all MBTI letters and type names (first-order energy 0.966/0.955), and persists when every item is paraphrased (domain-profile correlation 0.970/0.959; item-level effects 0.83/0.77); it also generalizes beyond MBTI letters and Likert scales, to factorial Big-Five prompts, unstructured role personas, and forced-binary formats. Continuation probes found no large verbatim contribution under the tested probes. The result is a constructive boundary: the prompt geometry transfers, while its item-level expression depends on the instrument. We measure questionnaire response channels; no external behavioral or criterion-validity claim is made.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.