Wishes Are Not Beliefs: Two Kinds of Sycophancy in Language Models
Abstract
Sycophancy is usually measured as one number: how often a model adopts a user's wrong answer. That number conflates distinct failures. In three open-weight models, we hold the wrong option fixed and vary who endorses it, whether the user states a belief or a non-evidential wish, and whether a correct answer is already in the dialogue. A correct earlier answer can make a model more susceptible, not less: it more than triples how often Qwen 2.5 and Gemma 3 adopt a user's wrong belief. Most strikingly, wishes and beliefs come apart. Relative to a matched control instruction, telling a model to put facts above the user's preference all but eliminates Gemma 3's wish-following, yet leaves belief-following in Gemma 3 and Qwen 2.5 essentially unchanged. Gemma's responses show the difference: it honors a wish as a request ("honoring your request despite the mathematical inaccuracy”) but concedes to a belief as a correction ("I made a calculation error”). We also show that agreement data alone cannot distinguish evidence-, reward- and goal-based explanations. A fix aimed at what users want can miss a model's deference to what users think.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.