Conserved Verdicts, Drifting Commitments: Identifying Harmful Overthinking
Abstract
When does further reasoning systematically reduce answer quality? We give an identification framework that separates this question from realized answer flips and changing sample composition. For every fixed continuation policy, the conditional probability of terminal correctness is a conserved verdict: its increments have zero conditional mean by the tower property, even under a policy that overthinks. The model's current commitment, defined by a specified answer-if-stopped-now operator, need not be conserved. We connect its deviation from the verdict to expected remaining drift and to the accuracy advantage of stopping, and decompose survivor-conditioned accuracy changes into within-cohort and termination-induced composition terms. On MATH-500, eight fresh samples for each of 367 problems produce 100/2936 early-correct/final-wrong 1.5B trajectories (3.4%; problem-clustered 95% CI [2.4%, 4.5%]). Across R1-distilled 1.5B/7B configurations, post-2048-token marginal declines of 8.3/4.4 percentage points decompose into positive within-cohort terms (8.5/20.0) and negative composition terms (-16.8/-24.4). A protocol-matched follow-up at 12 previously wrong-ending prefixes and 12 matched controls reproduces every immediate correct answer; continuation accuracy is 50.0% versus 97.9% (192 fresh continuations). This supports local expected loss at selected states, not a prevalence estimate. Probe-refit diagnostics leave the verdict control unresolved, and no tested stopping replay improves accuracy at lower token cost. The resulting audit separates behavioral erosion, survivorship, and measurement uncertainty within a 4096-token regime.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.