Judge2Vec: Selecting Compact LLM Juries from Unlabeled Judgments
Abstract
LLMs are increasingly used as judges because human evaluation does not scale, yet relying on a single judge can be unreliable and querying many judges is costly. We introduce Judge2Vec, an unsupervised framework that selects a compact jury from a large, heterogeneous pool using only the judges' multi-criterion ratings on unlabeled pilot examples. Rather than ranking judges by a single reliability score, Judge2Vec represents each judge with a criterion-specific competence embedding estimated through leave-one-family-out agreement with a family-balanced cross-family consensus, then composes a jury whose strengths cover different criteria while avoiding redundant judges. Remarkably, selecting fewer judges can improve both quality and efficiency: across eight text benchmarks spanning qualitative and quantitative evaluation, Judge2Vec achieves the highest performance among all compared methods, while four selected judges outperform full-pool aggregation and reduce average inference cost by 83.9%. Our approach also extends to multimodal and agentic evaluation. Beyond a fixed jury, the learned competence representations support budget-adaptive deployment and transfer across tasks, enabling jury construction for unseen tasks without any target-task examples or judge evaluations. These results show that LLM evaluation need not choose between a single unreliable judge and an expensive large jury: a small jury, intelligently composed from unlabeled judgments, can be both more reliable and more efficient.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.