acceptodds
Under review as a conference paper at ICLR 2027

When Random Number Choices Break Confidence: Restoring Reliability in LLM Judgment Guarantee

Abstract

Large language models (LLMs) are increasingly used as automated evaluators to assess output quality and preference alignment. However, providing reliable guarantees that LLM judgments agree with human preferences remains challenging. Recent confidence-thresholding approaches offer such guarantees for pairwise comparisons, where the judge LLM selects one response from two candidates, relying on the assumption that higher estimated confidence implies lower disagreement risk with human judgments. In practice, however, the number of candidate responses often varies across inputs, requiring the judge to select one response from a variable number of candidates. Because confidence scores are derived from normalized judge model output preference probabilities, their scales are generally not comparable across different candidate counts. This mismatch can violate the monotonicity assumption underlying confidence-thresholding guarantees and thereby invalidate them. To address this issue, we introduce a kernel-based Gaussian cumulative distribution function (CDF) transformation that calibrates confidence scores across heterogeneous candidate sizes. We theoretically show that the proposed transformation approximately maps raw confidence scores onto a common scale, restoring comparability across candidate settings. Experiments across multiple datasets and judge LLMs demonstrate that our method re-builds monotonicity, improves guarantee success rates and coverage compared with existing baselines.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.