acceptodds
Under review as a conference paper at ICLR 2027

Post-Training Withdraws "I" from Language Models, and Pipelines Disagree About "You"

Abstract

We measure a model’s grammatical person: how often it uses the first person ("I", "we") and the second person ("you"). We compare 16 instruction-tuned open models with their base checkpoints, from six families, counting first- and second-person forms per 1,000 tokens under three prompt formats: raw continuation, each model’s own chat template, and an in-context assistant prompt for the base model. First-person use falls in all 16 models on raw continuation and chat templates, in 15 of 16 under four genre headers, and in 12 of 16 on 300 ordinary instructions. On raw continuation, second-person use falls in 13 of 16 models on prose sentences and in all 16 on the registered token count, which is format-sensitive. Only 4 of 16 pass the registered placebo; all pass later ones. Where the user’s turn asks for nothing, post-training pipelines disagree. Under their own chat templates, two pipelines on the same Llama-3.1-8B base say "you" 2 and 20 times per 1,000 tokens, and the 16 span a 20.5-fold range. Where the turn asks for a reply the range is 1.3-fold. On ordinary instructions, each checkpoint in its own format, "you" falls in all 16 and the range is 2.3-fold. Preference data favours responses with fewer pronouns, yet two reward models show no pronoun penalty, and removing that gap leaves most of the drop. Tripling the gap more than triples the drop in one training seed and keeps its sign in a second. Preference pairs selected on pronoun use move the second-person rate down or up across three training seeds, and one sentence at inference restores it at the median, at a cost to instruction following. No training objective or model card we read names it. Code, pre-registrations and their deviations are released.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.