Correlation-Aware Ranking And Filtering For Heterogeneous Small-Model Judge Ensembles
Abstract
Shared errors limit the benefits of combining small language model judges. Judge CIA (Correlation-Informed Agents) constructs heterogeneous same-backbone LoRA pools, ranks adapters by validation accuracy, and filters pairwise error correlation before majority voting. Our analysis separates individual competence from effective jury size in a majority-error upper bound, identifying conditions under which a smaller, less correlated jury has a tighter bound. Across 12 RAGBench backbone–task settings, CIA improves accuracy over the validation-selected Top-1 adapter by 1.99 percentage points on average (paired-bootstrap 95% interval , conditional on fixed pools and selections), with gains up to 7.00 points. In four illustrative matched-size comparisons, CIA shows point-estimate test-accuracy gains of 2.81–4.74 points over accuracy-only selection, accompanied by lower test error correlation. For high-stakes judging, a slice-risk decomposition motivates weighted ranking. In a constructed compliance study, weighting reduces mean critical-example error from 2.26% to 1.65% and overall error from 2.60% to 2.48%. Public class weighting reduces target-class error by 8.17 points, with increased non-target error.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.