acceptodds
Under review as a conference paper at ICLR 2027

Diagnosing Harmful Continuation in Answer-Correct Mathematical Long-CoT Training Traces

Abstract

Long chain-of-thought (CoT) traces are widely used as supervision for reasoning-oriented LLM SFT, yet answer-correct traces can still lead to markedly different fine-tuning outcomes. We study post-conclusion continuation in answer-correct mathematical long-CoT data: a continuation where the answer appears sufficiently supported, but the trace continues with additional reasoning that remains in the supervised target. To test its training effect, we use a delete-only editor to construct answer-preserving suffix removal and compare CoT-based SFT on the original and processed traces. We observe improved SFT outcomes on mathematical problems after removing the editor-identified post-conclusion continuation, suggesting that this continuation is harmful to training in our setting. We therefore refer to this empirically supported phenomenon as harmful continuation. Beyond this intervention, we further characterize the removed post-conclusion continuation through uncertainty and hidden-state progress. We observe persistent local uncertainty together with weakened terminal-directional progress, forming an uncertainty–geometry mismatch. Finally, we introduce Harmful Continuation Cut (HCC), a lightweight proxy built on a frozen Qwen2.5-0.5B-Instruct backbone and trained to predict editor-identified boundaries with uncertainty–geometry auxiliary objectives. Our code is available at anonymous.4open.science/r/HCC-76B3.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.