When Do LLM Judge Panels Outperform a Single Large Judge? The Diversity-Bias Tradeoff
Abstract
LLM evaluation pipelines increasingly face a practical choice: use one strong judge or aggregate a panel of diverse judges. The usual argument for panels is variance reduction through less-correlated errors, but this view ignores an opposing effect: provider diversity can also increase bias heterogeneity. We formalize this diversity-bias tradeoff with a two-term decomposition of panel gain and derive a phase boundary characterizing when aggregation should outperform a single judge. Across four independently composed cross-provider panels spanning seven AI providers and four NLG benchmarks (21 dimensions, 2,635 items), provider diversity reduces inter-judge error correlation on every dimension tested (rho = 0.493-0.726 versus 0.925 for a same-provider baseline), yet lower correlation does not translate monotonically into higher evaluation accuracy because bias heterogeneity can offset the variance benefit. The decomposition predicts the direction of panel gain on 94% of dimensions satisfying its same-sign-bias assumption (81% overall). An equal-budget comparison finds a diverse panel approximately equivalent to a comparably priced single judge (10 versus 11 wins across 21 dimensions), showing that apparent panel advantages can reflect judge quality rather than diversity alone. Item-level analysis provides additional directional support for the variance-reduction mechanism (Spearman r = -0.10, p < 0.0001, N = 4,200). Finally, we translate the theory into a validation-based deployment rule that selects between panel and single-judge evaluation at the dimension level. These results explain when judge diversity helps, when heterogeneous bias cancels its benefit, and why panel deployment should be selective rather than automatic.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.