acceptodds
Under review as a conference paper at ICLR 2027

Does Assigned Identity Change How Language Models Use History? Separating Self-Binding from Focal-Agent Effects

Abstract

When a language model treats a history as its own, does that history receive privileged behavioral influence beyond the influence of a task-relevant non-self entity under matched prompts? We study this question with forced-choice tasks whose designated actions follow explicit simulator policies. A 320-item preregistered reviewer-control benchmark holds each history body fixed while crossing assigned identity with header order and adding an evaluation-only FOCAL condition. FOCAL is task-relevant but also changes the assistant to an outside-evaluator role, so SELF-FOCAL is an operational residual rather than a causal isolation of selfhood. Base Qwen 4B shows SELF-OTHER sensitivity +0.72 (95% CI [+0.57,+0.86]), with SELF-FOCAL +0.28 [+0.17,+0.41] and FOCAL-OTHER +0.43 [+0.25,+0.61]. After SELF/OTHER/NEUTRAL mixed training, SELF-OTHER is +0.73 [+0.58,+0.90] and its base-to-mixed change is +0.01 [-0.12,+0.15], while SELF-FOCAL changes by -0.33 [-0.54,-0.10] and FOCAL-OTHER by +0.35 [+0.05,+0.62]. Mistral 7B shows a different checkpoint-level pattern. Exploratory analyses of the frozen outputs show that the Qwen decomposition changes are sign-stable to leaving out any one training seed, but the study still has only five independent mixed-training runs. Current-facts and irrelevant-metadata controls narrow simple calibration explanations without identifying a unique mechanism. The methodological result is that SELF-OTHER alone is insufficient: role framing, focality/task relevance, header position, training exposure, and task-family composition must be separated before attributing behavioral privilege to self-binding.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.