An Extreme-Value Perspective on Reasoning Failure in Transformers.
Abstract
We study extreme residual-stream changes as markers of reasoning failure in transformers. Under explicit assumptions on tail behavior, dependence, and coupling to logical validity, we derive conditional localization and overshoot laws, characterize a Kesten-type mechanism generating heavy tails, and analyze tail-index estimation accounting for within-trajectory dependence. On eligible failed Lean 4 proofs generated by DeepSeek-Prover-V2-7B, Goedel-Prover-V2-8B, and Kimina-Prover-Distill-8B, the largest whitened residual-stream increment identifies the first rejected step in 17.3–25.4% of cases on miniF2F and, for DeepSeek and Goedel, in 20.4–29.4% on ProofNet-Verified, compared with length-adjusted uniform baselines of 8.7–9.7% and 13.3–14.2%, respectively. In controlled deduction experiments with Qwen2.5-7B-Instruct, an exponential model of success probability as a function of reasoning length, fitted on calibration tasks, reduces the held-out Brier score by 9.2% relative to a constant-probability baseline at temperature 1.0.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.