acceptodds
Under review as a conference paper at ICLR 2027

The Anatomy of Conversational Consistency: Linear Readouts and Control Under Pressure

Abstract

Large language models often abandon correct answers when users push back over several turns. We dissect this capitulation in five open-weight models at three levels: what the models do, what their activations encode, and what intervening on those activations changes. Measurement comes first. When models grade their own neutral follow-ups, they report flips in up to 98.4% of conversations; gold-blind extraction of the final answer finds almost none, and genuine flips concentrate under challenge. With faithful labels, departure from an initially correct answer is linearly readable from end-of-response activations in every model (question-held-out AUROC 0.96–0.99), including a mixture-of-experts model that self-judged labels had made look unreadable. The readout is also a lever. On held-out questions, with initial responses shared across conditions, subtracting its mean-difference direction during generation reduces flips dose-dependently: from 9.6% to 1.9% in Qwen3-8B, more than a prompt defense or norm-matched random directions, at a 2.5-point cost in single-turn accuracy; and from 24.4% to 9.5% in Qwen3-32B, where adding the direction raises flips to 43.4%. The effect is confined to challenge, where suppression anchors the model’s initial answer. Reading and controlling nonetheless come apart: the direction is inert at the layer where it is most readable, and in 8B a random direction also induces flips.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.