Flat by Theorem: What Truncated Chain-of-Thought Accuracy Curves Measure
Abstract
A growing literature asks when a chain of thought (CoT) decides its answer: truncate it, read out an answer, plot accuracy against the truncation point, read the rise as the moment of commitment. The quantity it measures, the probability that a trajectory ends correct given its prefix, is a bounded martingale under exact continuation, so its expectation is constant in the index and equals full-CoT accuracy; the companion agreement curve is a submartingale that rises for every model. Any rise is therefore an artefact; we decompose that of a published-style pipeline, rebuilt on our own traces, into three errors: a readout that intervenes rather than conditions, survivorship, and an index that is not a stopping time. On 8 models and 4 datasets (1,237,682 generations) the exact-continuation curve rises on no arm, equivalent to no change within 0.03–0.05; the forced-readout curve on the same traces climbs by +0.03 to +0.64 — by an amount that tracks the ordinary chain-of-thought gain closely enough to predict it: a rule registered on 6 arms before a 14B model was run put its effect at +0.520 [+0.396, +0.643]; it measured +0.633. Such a curve measures what the chain is worth, not when the answer was fixed. The readout dominates: survivorship alone manufactures declines of up to -0.464, but paired within a trace it and index leakage are null. Of 18 chat endpoints audited over one access path, 12 silently ignore assistant prefill, 10 in every sample; and a continuation conditioned on a re-encoded prefix passes every check a single-arm study can run, failing only the one that needs two. What survives is the corrected agreement curve's shape, free at half depth between 0.15 and 0.50. Code and records accompany this submission.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.