acceptodds
Under review as a conference paper at ICLR 2027

Ahead of the Evidence: Certifying Premature Commitment in Chains of Thought

Abstract

A chain of thought is increasingly read as evidence of how a model reached its answer. The worry that the model already knew before it finished reasoning is easy to state and hard to test, because "already" has no reference point: early relative to what? We supply one. On tasks whose correct procedure can be written as a short program, the program's state after each step fixes how far the answer is pinned down, a prescribed trajectory that no rule reading only that state can exceed. A chain that is really carrying out the procedure is one such reader, so it cannot have settled sooner. Earliness thereby becomes a testable claim rather than a suspicion. We locate a model's commitment causally, by transplanting the hidden state from one step of a different problem into the chain and asking whether the answer follows the transplant. Causal Lookahead Probing (CLP) fits no parameters, so a positive reading cannot be the artifact of a learned mapping that has proved fatal to probe-based causal analyses. On transformers compiled from the program itself, where the computation is known in advance, the method recovers that trajectory to within at every step. Across four such tasks and nine checkpoints up to 8B, reasoning-distilled models commit up to earlier than that state could license on three of them; on the task that licenses commitment latest, the gap to the undistilled parent persists at equal accuracy. Wherever the bound has room to bind, their undistilled parents stay within it, as does every base, instruction-tuned and state-space control. Their chains are, in these cases, transcripts of a decision already made.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.