acceptodds
Under review as a conference paper at ICLR 2027

AURA: Adaptive Uncertainty-aware Refinement for LLM-as-a-Judge Auditing

Abstract

Large language models (LLMs) are widely used as judges for open-ended generation because human evaluation is costly and hard to scale, yet their preferences remain imperfect proxies for human judgment. Human verification can reveal which judge decisions to trust, but it is typically affordable for only a small fraction of them. We propose AURA, an adaptive, uncertainty-aware framework that jointly refines estimates of judge–human agreement and guides human verification for pairwise LLM-as-a-judge decisions. The framework updates a probability of agreement for each comparison using a learned human-consistency scorer and local and anchor-based evidence. Sparse transport propagates reliable signals to uncertain comparisons, while uncertainty guides the allocation of additional human verification. Our synthetic and real-data evaluations show how performance varies with judge quality and verification budget, and illustrate when adaptive refinement can help allocate limited human review.

open until 14 Dec 2026

est. 32% chance this paper gets accepted at ICLR 2027.

Reject 68%Accept 32%

What do you think this paper will get?

All positions stay anonymous.

Related papers

Loading the map…

Discussion (0)

Sign in to comment.