acceptodds
Under review as a conference paper at ICLR 2027

OPENING LLM JUDGES: RECOVERING PREFERENCE SIGNALS BEYOND THE FINAL VERDICT

Abstract

LLM judges are widely used to evaluate model outputs, but their verdicts can be unreliable: a judge may favor the worse answer because of its position, length, or other surface features. When a judge gives the wrong verdict, is the information needed to make the right judgment absent from the model, or is it still present in its internal representations but not reflected in the final output? We study this question across 64 open-weight evaluators and 14 datasets, including causal interventions on 41 judges (editing a model's internal activations while it runs to see whether its verdict changes). On LLMBar, a benchmark built so that the superficially better answer is the worse one, the verdicts of 50 judges agree with human labels only 0.456 of the time, even after position bias is cancelled by scoring both answer orders. Yet a small probe trained on the same judges' internal activations, without changing the judges, reaches 0.846, and still 0.686 after the influence of surface features such as length and position is removed (0.507 with shuffled labels). This gap holds across eight benchmarks and across model families, but it is not universal. A simple score of how well surface features alone predict the human label, computed before any probe is trained, is strongly correlated with the size of the gain (Spearman rho = 0.90). On two further tasks where a judge scores one answer at a time against a rubric, leaving no surface cue to exploit, reading the internals gives no advantage. The interventions also show that editing activations in the middle of the network already changes the verdict, before it can be read off directly, and locate the pathways that carry position and length bias. With the same human labels, the recovered signal also lets a judge flag cases where it is likely to be wrong and yields better labels for preference learning. A wrong verdict, then, does not mean the judge lacks the information, and a simple diagnostic shows when it is worth recovering.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.