acceptodds
Under review as a conference paper at ICLR 2027

Committed by the First Token: Measuring Recovery from Induced Answer Commitment in Chain-of-Thought Reasoning

Abstract

Chain-of-thought is usually read as the process by which a model works out its answer. We ask what a chain does with an answer that is already present when it starts. We prefill the assistant turn with graded yes/no commitments and, by holding the reasoning chain fixed while varying what sits in context, decompose the resulting shift in the elicited answer distribution into a direct path, through which the commitment is read at answer time, and a mediated path, carried by the chain the commitment induced. On BoolQ with Qwen3-4B a plain commitment overturns about half of the correct answers it contradicts, and the mediated path dominates: a chain written under a plain “no”, with the commitment itself removed, displaces the answer by 10.9 logits, against 5.0 for the commitment alone. Step-level probing shows the hand-off directly: the direct pull decays over the chain while the mediated pull grows, the two crossing within five steps. Three interventions confirm the split: masking the commitment from attention during generation drives the mediated effect to zero; withholding the supporting passage leaves the direct path unchanged and roughly doubles the mediated one, so the chain is not carrying the commitment by quoting evidence for it; and a chain generated under the opposite commitment overturns the commitment still beside it, driving compliance from 0.85 to 0.24. The model cannot report an unmarked commitment as foreign, and being told does not reduce compliance; directing the chain does. A forced verification prefix lowers compliance with a wrong commitment in every condition, and attenuating attention to the commitment to a tenth of its weight recovers 79% of the net damage while keeping the commitment readable. On Ministral-8B the effect is larger, not smaller, and the commitment's strength—legible before reasoning begins—no longer orders the outcome by the end of the chain.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.