Predicting Reasoning Errors Before They Occur: A Prompt-Side Activation Direction Anticipates Correctness in a Large Language Model
Abstract
Responses produced by large language models are eloquent in a way that makes right answers indistinguishable from wrong ones. It is critical to predict the correctness of the model’s response, which significantly helps in either terminating the model from answering or routing it to a bigger model instead of producing incorrect answers eloquently. In this study, we use activation patching to identify a candidate layer for the arithmetic reasoning task. We utilize the hidden state activations at the candidate layer as a direction to develop a difference-in-means measure. We test the performance of the approach across three models: Mistral-7 B, Llama-2-13B, and Qwen 27B. The measure predicts subsequent correctness across all three models (AUC .627, .673, and .786 from smallest to largest model), and the signal emerges around the middle of the network and strengthens toward the task-relevant layers. The learned direction transfers to new formulations and an independent benchmark without retraining (AUC .810 and .864 on the largest model), although its decision threshold shifts across datasets and requires recalibration. These results place correctness information earlier in the computation, before the model commits to an answer, suggesting that reliability can be assessed before generation rather than only diagnosed after it. The principal methodological contribution of this paper is to utilize the activations of causally identified candidate layers and develop a lightweight difference-in-means measure to predict the correctness of models’ responses before the generation of answer tokens.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.