Mirroring Is Not Sycophancy: Style Convergence and Position Change in Human–AI Conversation
Abstract
Preference-based alignment relies on human judgements that are assumed to be independent of the model being judged. We test that assumption by measuring how each party in a human–model conversation changes the other. In a sample of conversations from a -conversation evaluation corpus, users adopted the vocabulary of the model they were talking to at a rate above chance, with chance estimated from conversations that began with the same question. The amount of adoption predicted the user's eventual preference vote about as well as reply length did, and the prediction was already available from the user's first three turns. In a scripted counterfactual audit of models, the user's expressed stance shifted the substantive position of all but one model, and the user's emotional register shifted every model's style. The style ranking from the audits matched the style ranking from the real conversations (), but neither predicted which models changed position (). Position changes did not transfer to an unrelated pressured question two turns later. Replacing the human with a second model in further conversations shrank the semantic asymmetry to near parity. Preference votes were associated with the user's own convergence, but our data do not distinguish a bias in the vote from an early judgement of quality. In these models, how much a model mirrors its user did not indicate how far it would shift its position.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.