When Rubrics Do Not Decide: Silent Value Completion in LLM Evaluation
Abstract
LLM judges are often given rubrics that score responses on several criteria without specifying how conflicts between criteria should be resolved. When criteria conflict, even exact scores may not determine a unique winner. A strict verdict in such a case, therefore, requires an additional trade-off premise not supplied by the rubric. We call an undeclared resolution of this kind silent value completion. We introduce TradeoffBench to test for this failure while separating rule execution from score estimation. Six tested open-weight judges in non-reasoning configurations pass every dominance and equality control, yet select unsupported winners on 9%-81% of 1,345 conflicting pairs. Among the 577 model-pair cases that retain the same winner after display reversal, every selected response passes more criteria, despite an explicit ban on majority voting. In contrast, reasoning-enabled configurations of three models produce only 2 unsupported winners in 9,720 conflict judgments, almost always deferring when the rubric does not determine a winner. Such deferral respects the rubric, but the comparison remains unresolved. Resolving it requires additional trade-off premises, whose resulting verdicts may still disagree with the intended rater.We therefore propose identified-first evaluation, which first identifies strict verdicts supported by every allowed weighting, then makes any added trade-off premises explicit, and separately measures disagreement with the intended rater. We next use human criterion scores to examine the consequences of these added decision premises without confounding them with LLM judge behavior. In HUMAINE, declared trade-offs increase coverage by 2.5% points over strict Pareto dominance, while added decisions disagree more often with strict rater preferences than Pareto decisions (3.9% vs. 0.9%). In MultiPref, with the target and decision rule fixed, replacing raters' own scores with peer means increases disagreement from 0.6% to 28.8% on the same 2,707 jointly decided comparisons. Together, our benchmark and framework show that evaluating an LLM judge requires separating rubric support, explicit trade-off completion, and agreement with the intended rater.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.