acceptodds
Under review as a conference paper at ICLR 2027

When Conclusions Reverse: Signed vs. Magnitude Effects in Chain-of-Thought Evaluation

Abstract

As reasoning models are increasingly evaluated by their chain of thought (CoT), perturbation tests are a common way to probe how much the chain affects the model’s answer by changing part of it and measuring how the answer shifts. The resulting effect can be scored in two ways: a signed score that keeps both the size and the direction of the change in the reference answer’s probability, or a magnitude-only score that keeps only the size. We ask whether this scoring choice alone can change the conclusions drawn from a CoT evaluation when everything else is held fixed. Across about 79,000 perturbation records from nine models and six question-answering datasets, we find that it can. (1) In our main comparison, the two scores rank two models in the opposite order over the 132 questions they share (out of 144 and 150); and (2) they rank a different part of the reasoning first for 40–56% of questions in three of four model–dataset combinations spanning two comparisons, one on a multiple-choice dataset and one on a yes/no dataset. In the broader descriptive analysis, model rankings change in most comparisons, while others show little or no disagreement between the two scores. These results show that signed and magnitude-only scores answer different questions and should be reported separately when both are relevant.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.