Evidence, Pressure, or Neither: What Actually Moves Costly Decisions in Large Language Models
Abstract
Language models are increasingly asked to commit to costly judgements and to revise them when challenged. Evaluations report a single revision rate against a control with the pushback removed, conflating revision that tracks the argument with revision that tracks the arguer and assuming that the control does not move. We propose Matched-Presence Decomposition (MPD), which separates revision into a content, a social and a re-ask component by replacing each manipulated element of the prompt with the same element from another item, together with six audits that decide whether each component is identified. On a prediction market, a bare second ask in the same conversation changed the answer in to of episodes in four open-weight families, and neither the framing sentence, greedy decoding nor dropping the word “updated” removed most of it. Across seven families, argument content moved every one and direction-free reputation none. The audits also mark where a reading fails: a peer-review placebo carries direction, and on political items part of what a published sycophancy benchmark scores is agreement with the persona's last assertion. A prompt that makes the model state whether the material is new removed to points of drift while preserving the response to evidence; the largest reduction, which a reduction-only evaluation would select, came from an instruction that removed up to of that response.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.