Hypothesis Anchoring: Evidence Use Without Belief Revision on Self-Generated Hypotheses
Abstract
Reasoning models often commit early to an incorrect hypothesis and never revise it, even when their own chain of thought contains the refuting evidence, and even when the correct answer is inserted into their context in plain text. We quantify this failure, which we term *hypothesis anchoring*, with an injection stress test on incorrect GSM8K traces of two open-weight reasoning models, replicated under two annotation protocols for locating the anchored hypothesis. Stating the ground-truth answer inside the model's ongoing reasoning barely increases the rate at which it reaches the correct answer, and a wrong-value injection does not reproduce the gain: insertions act through their content rather than their mere presence, and the model integrates true evidence only weakly. Re-solving the same problems from scratch does markedly better, so the anchored context itself is costly. In roughly a quarter of injected continuations the correct answer appears verbatim in the generated text, yet the model ends its reasoning elsewhere. We then ask where the anchor resides, excluding each candidate with matched controls and reported minimum detectable effects: we find no evidence that anchoring stems from inattention to counter-evidence (it is attended), from evidence unavailability, from any individual attention head or a single representational direction; deleting the hypothesis sentence does no more than a sham deletion, and blocking post-evidence attention to it, with fluency preserved, changes nothing. These nulls are stable across both annotation protocols and both models. A layer-wise attention signature behaves unstably across trace sources and is not advanced as a mechanism. Commitment survives every intervention we test within the anchored context, leaving training-time intervention and scale as the most plausible remaining levers. We release the protocol, annotations under both protocols, and code.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.