When Conclusions Reverse: Signed vs. Magnitude Effects in Chain-of-Thought Evaluation
Abstract
As reasoning models are increasingly evaluated by their chain of thought (CoT), perturbation tests are a common way to probe how much the chain affects the model’s answer by changing part of it and measuring how the answer shifts. The resulting effect can be scored in two ways: a signed score that keeps both the size and the direction of the change in the reference answer’s probability, or a magnitude-only score that keeps only the size. We ask whether this scoring choice alone can change the conclusions drawn from a CoT evaluation when everything else is held fixed. Across about 79,000 perturbation records from nine models and six question-answering datasets, we find that it can. (1) In our main comparison, the two scores rank two models in the opposite order over the 132 questions they share (out of 144 and 150); and (2) they rank a different part of the reasoning first for 40–56% of questions in three of four model–dataset combinations spanning two comparisons, one on a multiple-choice dataset and one on a yes/no dataset. In the broader descriptive analysis, model rankings change in most comparisons, while others show little or no disagreement between the two scores. These results show that signed and magnitude-only scores answer different questions and should be reported separately when both are relevant.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.