Distilled Reasoning Models Bypass Uncorrected Errors in Their Own Chains of Thought
Abstract
Reasoning models are increasingly built by distillation, that is supervised fine-tuning on chains of thought sampled from an RL-trained teacher. We study chain-of-thought faithfulness in these models through a regime we call bypass, in which the model produces the correct final answer while an uncorrected error stands in its chain of thought, and we ask how often it occurs on the models' own errors, what governs its rate, where the answer then comes from, and what training changes it. We measure bypass by injecting a wrong step mid-thought and regenerating, and by auditing the models' own arithmetic errors across roughly thirty thousand sampled rollouts, executing every equation they write. Both instruments are checked against 492 blind human labels. On its own validated errors the 7B distill bypasses a quarter to a third of them depending on the detector, with blind human annotation corroborating the effect, and two thirds on a memorization-proof synthetic set. Under injection, bypass reaches half to three-quarters of injected errors on GSM8K and on the synthetic set across the DeepSeek-R1-Distill-Qwen family under the judge, whereas the base-matched RL model QwQ-32B shows near zero on GSM8K. It falls to roughly 3% on AIME and is several times lower on BBH state-tracking, a pattern consistent with answer recomputability across our benchmarks, and scale does not remediate it. Attention knockout, forced-answer, and interruption probes locate the correct answer in a post-boundary re-derivation with no detectable direct influence of the corrupted step, and injecting at a fixed early position overstates bypass for the distilled model by fourteen points relative to its own errors on the same items, a gap that closes when the injection is placed at the natural error position. RL post-training on the same distilled checkpoint cuts bypass three- to four-fold under injection and by a third to a half on the model's own errors. Bypass marks a regime in which the visible chain ends in the wrong state while the correct answer is re-derived after the reasoning boundary, so a monitor reading the chain sees wrong reasoning and a right answer with no correction between them, and RL post-training of the same checkpoint reduces it substantially.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.