acceptodds
Under review as a conference paper at ICLR 2027

Do Injected Errors Transport to Organic Errors? A Matched Audit of Corruption Probes for Chain-of-Thought Reasoning

Abstract

Results from injected-error probes do not transport to the errors language models actually make at the rates a probe implies. We audit the standard corruption-probe protocol on frozen frontier models, holding the item, the prefix position and the trace author fixed. In a pre-registered pilot (three proprietary families, three public math benchmarks, four typed corruptions, 2,976 continuations), injected corruptions were recovered from without any acknowledgment in 37.8-60.2% of continuations per family (recovery of any kind 86.8-97.2%; invisible step-deletions excluded and reported separately), while the models' own reproduced errors were repaired in 0 of 60 error-localised continuations. The author-fixed contrast is i-c: on the same items and positions, a corruption injected into another model's correct trace is repaired far more often than that same model's real error; authorship is fixed, but the error and its preceding derivation both differ. In the exploratory cohort the paired gap is +0.40 to +0.75, every Newcombe 95% interval excludes zero, and exact McNemar tests Holm-adjusted jointly over the six primary contrasts are significant at 0.05 in three and reach adjusted p = 0.0645 in the other three (discordant pairs 6/0, 6/0, 9/1). A separately registered confirmatory cohort on untouched items contains 173 b-c matched pairs and an eligible subset of 122 i-c matched pairs across 1,038 continuations. All six primary point estimates are positive, matching the exploratory signs, with author-fixed point-estimate gaps of +0.227 to +0.415; after Holm correction only the two gpt contrasts reach 0.05, while the sonnet adjusted values are 0.0509 and 0.068 and flash is not separated from zero. The confirmatory point estimates therefore retain the exploratory direction, but this study does not establish that direction separately in every family after multiplicity correction. Two further exploratory arms probe the error distribution itself: transplanting the real error TEXT into a correct derivation gives j-c point estimates of +0.00 to +0.06 with every interval spanning zero, which neither establishes equivalence nor identifies a mechanism, while a formulaic corruption at the same position in the same trace is repaired far more (i-j +0.39 to +0.75), and cutting just before the error restores recovery (f-c +0.39 to +0.70). Exploratory mechanism probes are consistent with a surface-form account: a blind judge flags 66.1% of injected steps as artificially edited versus 20.6% of real error steps, while a small reference LM's token surprisal does not separate the two (AUC 0.47, null). The higher edit-flagging rate is consistent with a surface-form hypothesis, but the reference-LM result is inconclusive and neither probe identifies the cause of the recovery gap. Transplanted real errors are flagged at 31.7% against 20.6% for real errors, an intermediate rate this sample does not resolve; their j-c recovery estimates run from +0.000 to +0.056 with every 95% interval including zero, which establishes neither equivalence nor the absence of a recovery advantage. Recovery estimates from these injected-error probes should therefore not be substituted for organic-error recovery rates without validation on the target error population, and outcome-label calibration needs recovery rates from that population. All corpora, continuations, the validated acknowledgment detector and registration artifacts accompany the submission.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.