The Evaluation Bottleneck: Your Biggest Judge Is Not Your Best Judge
Abstract
Agentic LLM systems generate outputs faster than they can be reliably reviewed: generation is cheap, evaluation is not. The common responses are to make the judge bigger or more specialized. We ask a more basic architectural question: where should specialization live in an evaluation system—in the judge itself, or in the rule that decides when its verdict can be trusted? We test both placements. In the first, we split a judge's training data across expert evaluators, one per family of criteria. In the second, we keep judges of different sizes frozen and train a small trust head that either accepts a cheap judge's verdict or defers to a stronger one, with every automatic release audited against a predeclared risk target. The findings reverse intuition. Specializing the judge backfires: splitting supervision across eight experts costs 10 accuracy points and shrinks the fraction of judgments that can be safely released more than fourfold—a loss that vanishes once the experts start from a shared trained judge. Specializing the trust decision wins: on RewardBench 2, a 15-parameter trust head routes a cascade of small-to-large judges to an accuracy 4.7 points above the largest judge alone at 0.42× its compute, passing exact 95% risk audits in all 20 runs, while naive confidence-based escalation yields neither an accuracy gain nor compute savings. The design rule is simple: share judgment learning by default, and specialize when to trust, not what judges know—provided the earlier judges are genuinely cheap.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.