Metric Match: A Subset Selection Approach to Evaluating LLM Judge Reliability
Abstract
LLM judges are used to reduce the need for costly human labor in evaluating open-ended text generation. However, the reliability of these judges depends critically on their alignment with human raters — a property that itself depends on costly human annotations. In this work, we develop a method (Metric Match) for estimating correlation-based reliability metrics of LLM judges from limited annotations. Metric Match selects a subset of samples for human annotation such that the subset matches the population reliability metric with respect to acquired synthetic labels. We empirically show that Metric Match achieves a win-rate of at least 0.794 with respect to four natural baselines across four different correlation metrics and 15 datasets. Furthermore, Metric Match exhibits a 16.2% decrease in average estimation error and reduces annotation needs by 28.1%. We provide a cost model and highlight a medical case study where our method saves $1,083.33 on expert annotations compared to baselines. Finally, for the downstream task of classifying whether a judge is sufficiently reliable, Metric Match also outperforms all baselines in terms of accuracy.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.