acceptodds
Under review as a conference paper at ICLR 2027

Right Verdict, Wrong Reason: Counterfactual Attribution Reveals Shortcut Learning in LLM Evaluators

Abstract

As large language models take over the role of evaluator—scoring outputs, modeling rewards, and checking policy compliance—trustworthy evaluation is emerging as a problem as hard as capability itself: an evaluator's verdicts gate agent actions, human review, and training data, yet the evaluator is validated only by the labels it emits. Deployed evaluators are, moreover, continually updated: evidence pipelines are corrected, rules are amended, and the scope of a rule is widened or narrowed. When a verdict flips after an update, standard evaluation checks only whether the new label is correct—never whether it changed for the right reason. This failure mode, right verdict, wrong reason, is invisible to accuracy and is precisely what auditability requires ruling out. We make it measurable. Factoring an evaluator's inputs into typed components (evidence, rules, scope), the 2^3 old/new hybrid configurations form a counterfactual table whose inclusion-minimal component swaps define an exactly computable gold attribution for any verdict change; we prove correctness of enumeration, bound attribution multiplicity, establish a worst-case query lower bound for black-box verification, and separate predicting an attribution from verifying one. We instantiate the framework in ReasonBench, a benchmark of 19,520 cases with executable rule-based evaluators, exact gold attributions, and paired consistency controls. The findings are cautionary. Fine-tuned models reach 98.4% exact attribution accuracy in distribution, yet merely reordering semantically equivalent input sections changes their attribution about half the time; under compositional shift, verdict accuracy holds at 93.8% while attribution accuracy on changed cases collapses to 7.2%; and a pre-specified hypothesis that supervising the full counterfactual table would improve attribution is rejected—the richer target consistently hurts. Accuracy is a poor proxy for faithful attribution; reliable attribution requires executing counterfactuals rather than trusting model outputs.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.