acceptodds
Under review as a conference paper at ICLR 2027

SmallJudge-Bench: Evaluating Correct Invariance in Evidence-Grounded Judges

Abstract

An evidence-grounded judge should retain a correct verdict when a claim is rephrased without changing its factual content or relation to the supplied evidence. Prediction stability alone misses this requirement because a judge can be consistently wrong. We introduce SmallJudge-Bench, with 2,800 individually reviewed original–edited pairs across four datasets. Claude Sonnet 4.6 exceeds a clean-trained ModernBERT-large encoder in paired correctness by 4.80 percentage points (95% confidence interval ). Yet, on originals both models answer correctly, Sonnet's edited error rate is 2.64 points higher (). Matched subsets from two editing generators show the same conditional ordering, although the interaction between generator and model remains unresolved. Label-preserving training augmentation improves paired correctness and reduces prediction changes relative to repeated-clean training. These findings distinguish overall correctness from retaining correct decisions under edits, within the evaluated configurations. The protocol reports both, with separate controls for evidence-free stability and responses to changed content.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.