Is Sycophancy a Form of In-Context Updating?
Abstract
Factual sycophancy, a language model's deference to a user's false beliefs, is usually understood as a bid for the user's approval at the expense of truthfulness. We argue that this framing over-emphasizes the user and propose to view sycophancy as a failure within the broader and generally useful process of in-context updating, by which models revise their answers based on information supplied during an interaction. We support this view with behavioral and mechanistic evidence. First, models' responses to user claims sit within a broader pattern of source-sensitive updating. They revise their answers in response to claims attributed to other sources as well, and the effect scales with how well-supported the claim appears. Second, the internal components that support deference to the user are also causally important for updating from other sources. Together, our results suggest that factual sycophancy is a calibration failure of in-context updating, in which a user's assertion exerts more influence than its evidential value justifies. Because sycophancy is embedded in this process, reducing it risks reducing useful updating as well, which calls for sycophancy evaluations and mitigations that are sensitive to this trade-off.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.