acceptodds
Under review as a conference paper at ICLR 2027

Introspective Verification: Reading Step Correctness at Intermediate Depth

Abstract

Process reward models score the correctness of individual reasoning steps, and these scores drive selection and search in multi step reasoning. The most accurate of these models work through generation, decoding a critique for every step and often calling an interpreter, which makes checking a solution expensive. Whether the same judgment can instead be read directly from the verifier's own hidden states has not been resolved. We use a linear probe that locates the signal at an intermediate layer near seventy percent of the depth, where step correctness is more separable than at the final layer. We call this introspection, since the verdict is decoded from the backbone's own representation rather than generated as text. Once each readout is trained rather than probed, however, the final layer becomes the stronger readout on its own. Our introspective verdict head sums a trained intermediate readout and a trained final readout into a per step reward in a single forward pass, with no generation and no code execution. We adapt the backbone with LoRA on about eight thousand conversations from one released label set, and both training and inference run on a single 32GB V100 at the 1.5B and the 7B scale. At 1.5B the fusion hyperparameters are tuned only on MATH and applied unchanged to three domains never touched during tuning, which shows cross-domain generalization. On ProcessBench the 1.5B head reaches 58.1 mean F1, above most trained process reward models including ones at 7B and 14B scale, at about one quarter the token cost of a single generative pass. Fused with a 7B generative verifier it raises mean F1 from 75.2 to 77.4, without lowering any subset. Code is available at https://anonymous.4open.science/r/LLM_CALM-48B7.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.