acceptodds
Under review as a conference paper at ICLR 2027

Mirroring Is Not Sycophancy: Style Convergence and Position Change in Human–AI Conversation

Abstract

Preference-based alignment relies on human judgements that are assumed to be independent of the model being judged. We test that assumption by measuring how each party in a human–model conversation changes the other. In a sample of conversations from a -conversation evaluation corpus, users adopted the vocabulary of the model they were talking to at a rate above chance, with chance estimated from conversations that began with the same question. The amount of adoption predicted the user's eventual preference vote about as well as reply length did, and the prediction was already available from the user's first three turns. In a scripted counterfactual audit of models, the user's expressed stance shifted the substantive position of all but one model, and the user's emotional register shifted every model's style. The style ranking from the audits matched the style ranking from the real conversations (), but neither predicted which models changed position (). Position changes did not transfer to an unrelated pressured question two turns later. Replacing the human with a second model in further conversations shrank the semantic asymmetry to near parity. Preference votes were associated with the user's own convergence, but our data do not distinguish a bias in the vote from an early judgement of quality. In these models, how much a model mirrors its user did not indicate how far it would shift its position.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.