acceptodds
Under review as a conference paper at ICLR 2027

DisJudge: Buying Reliable Adjudication Where Candidates Disagree

Abstract

Large language models often write several candidate programs for one task, and a common way to select one of them is to count the generated tests that each candidate passes. The expected outputs of these tests come from an oracle, which is usually a model and often the one that wrote the candidates. We show that the count compares two candidates only on the inputs where they disagree, so only the accuracy of the oracle on this disagreement set matters, and its overall accuracy can hide a strong bias there. On six code benchmarks, a same-model oracle is right on of all inputs, yet on the disagreement set it returns the output of the correct candidate of the time and that of a wrong candidate of the time. We prove that for a pair of candidates the count selects the wrong one with probability tending to one whenever the oracle favors it on this set, whereas a calibrated likelihood score stays consistent at a rate set by the Hellinger distance between the answers of the oracle under the two hypotheses. Based on this analysis we propose DisJudge, which spends a small budget of reliable answers on the disagreement set and ranks the candidates by the calibrated likelihood of all answers. With reliable calls per disagreement input, DisJudge raises the mean selection accuracy over benchmark and generator settings from to , while the best of baseline rules reaches . The bias, and with it the gain of DisJudge, shrinks as the oracle shares less with the generator.

Then back it, or bet against it.

Related papers

Open the market on this paper to see 7 more related papers.