Certifying No-Winner Decisions in Large Language Model Judging
Abstract
Large language models (LLMs) are increasingly used as automatic judges of response quality and preference alignment. Existing guarantee-based evaluation frameworks typically assume that each instance has a unique human-preferred response, an assumption that often fails in open-ended settings where candidates may be similarly poor leaving no meaningful winner. We study reliable LLM judgment under this more general setting. Allowing no-winner outcomes introduces a key challenge: confidence scores can conflate evidence about absolute response quality, relative preference, and candidate ambiguity, complicating the identification of reliable no-winner decisions. Moreover, controlling overall selective risk does not ensure reliability separately for no-winner and response-selection decisions. We address this challenge with a two-stage certified selection framework. First, conformal prediction constructs a localized label set over candidate responses and the no-winner label, retaining the reference label with a distribution-free, finite-sample marginal coverage guarantee under exchangeability. Second, conditioned on the localized set, separate calibrated hypothesis tests certify either a no-winner or response-selection decision. The resulting framework provides high-probability selective correctness guarantees while abstaining or escalating when evidence is insufficient. Extensive experiments demonstrate the benefits of our two stage framework for no-winner recognition and characterize its effectiveness across controlled and human-rated evaluation settings.
Then back it, or bet against it.
Related papers
Open the market on this paper to see 7 more related papers.