Reasoning Change-Points (RCP): Training-Free First-Error Localization from Hidden-State Dynamics
Abstract
Hidden-state representations provide a rich substrate for diagnosing multi-step reasoning, yet when reasoning failures leave a reliably localizable boundary in representations remains insufficiently understood. Our key insight is to separate the location of a reasoning error from the representational change that makes that error detectable. We find that boundary strength varies systematically across reasoning regimes: controlled perturbations often induce sharp and reproducible transitions, while naturally occurring failures exhibit subtler structure under passive inspection. Importantly, this structure is not intrinsic to the failure alone: re-evaluating reasoning prefixes under an explicit verification instruction substantially strengthens natural localization, showing that first-error localizability depends not only on the reasoning failure itself but also on the computation used to inspect it. Building on this observation, we introduce Reasoning Change-Points (RCP), a reliability-aware framework that detects local transitions in step-level hidden states, estimates the reliability of candidate error boundaries, and strengthens ambiguous natural evidence through verification conditioning, without training an auxiliary verifier. Across six 4B–8B models spanning multiple model families on GSM8K and MATH-500, controlled failures achieve 0.51–0.90 exact localization, and a prospectively specified Qwen3-32B evaluation reproduces the predicted depth structure. Across four ProcessBench domains, verification conditioning improves exact natural localization by 17.5–22.8 percentage points over passive hidden-state trajectories, indicating confidence calibrated on controlled failures transfers to natural settings, substantially enriching for correctly localized examples. Downstream experiments show that informative boundaries can support targeted regeneration and motivate adaptive reasoning control. Together, these findings characterize a soft reliability boundary for hidden-state first-error localization and show how that boundary can be strengthened and selectively exploited when the resulting evidence is reliable.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.