SmallJudge-Bench: Evaluating Correct Invariance in Evidence-Grounded Judges
Abstract
An evidence-grounded judge should retain a correct verdict when a claim is rephrased without changing its factual content or relation to the supplied evidence. Prediction stability alone misses this requirement because a judge can be consistently wrong. We introduce SmallJudge-Bench, with 2,800 individually reviewed original–edited pairs across four datasets. Claude Sonnet 4.6 exceeds a clean-trained ModernBERT-large encoder in paired correctness by 4.80 percentage points (95% confidence interval ). Yet, on originals both models answer correctly, Sonnet's edited error rate is 2.64 points higher (). Matched subsets from two editing generators show the same conditional ordering, although the interaction between generator and model remains unresolved. Label-preserving training augmentation improves paired correctness and reduces prediction changes relative to repeated-clean training. These findings distinguish overall correctness from retaining correct decisions under edits, within the evaluated configurations. The protocol reports both, with separate controls for evidence-free stability and responses to changed content.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.