If You Know, You Know: Self-Supervised Uncertainty Estimation for LLM Judges Over Binary Choices
Abstract
'LLM-as-a-judge' systems are increasingly relied upon for numerous tasks, such as labeling training data for AI models, and evaluating LLM generations in industry applications. However, a pervasive problem of these systems is that LLMs exhibit an alarming tendency to reverse their preference over a binary choice, when the presentation order of the choice is inverted. This behavior is patently irrational, and calls into question whether it makes sense for the field to rely on LLM judges, absent a reliable measure of uncertainty over their verdicts. To that end, we propose a simple, scalable, and self-supervised method of training LLM judges to measure their own uncertainty in preferences over binary choices. We show that our method is more accurate than any of our tested baselines, and generalizes surprisingly well across data domains. We also show that our method scales gracefully across model-sizes, and transfers across model families. Finally, we show that our method can enable downstream use-cases, such as adaptive inference compute for more accurate LLM-as-a-judge pipelines.
est. 32% chance this paper gets accepted at ICLR 2027.
What do you think this paper will get?
All positions stay anonymous.