acceptodds
Under review as a conference paper at ICLR 2027

Quantifying Consensus: Overcoming Exact Match Limitations in LLM-as-a-Judge Evaluation

Abstract

As Large Language Models (LLMs) increasingly serve as autonomous evaluators for complex outputs (LLM-as-a-Judge), the demand for fine-grained evaluation on granular scales (e.g., 1-5 or 1-10) has grown. Yet, this paradigm shift has exposed a critical vulnerability: we are evaluating the alignment of the next-generation artificial judges using legacy metrics that were never designed for this task. Current evaluation methods primarily rely on rank correlation coefficients (e.g., Spearman’s ) or traditional Inter-Annotator Agreement (IAA) metrics like Cohen’s Kappa. Both approaches are fundamentally limited for modern LLM evaluation: rank metrics ignore absolute calibration, while standard IAA metrics demand exact numerical consensus and are highly susceptible to statistical deflation under skewed distributions. To address these gaps, we introduce Tolerance-Weighted Agreement (TWA, ), an absolute alignment metric natively designed for the LLM-as-a-Judge paradigm. By using graded tolerance weights to absorb minor subjective variance, TWA isolates true semantic consensus between artificial judges and human evaluators without suffering from statistical deflation. We validate the metric across an In-Context Learning (ICL) classification explainability dataset and the evaluation sets of established benchmarks (BiGGen-Bench, SummEval). The results demonstrate TWA's robustness in quantifying Human-LLM alignment across diverse tasks. Furthermore, it provides a highly interpretable alternative to complex metrics like Gwet’s AC2 and naturally scales to measure Inter-Judge Agreement directly between competing LLMs.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.